Overview In this role, you will push the performance of NVIDIA GPUs through deep analysis and optimization of core algorithms and data structures. You'll work across hardware and software teams to guide accelerated computing initiatives and co-design software with hardware. You will explore applications that stretch GPU architectures and contribute knowledge via scholarly and professional publications. This position offers impact at scale in a fast-moving, technology-driven environment and a path toward advancing cutting-edge HPC and AI workloads.
Compensation / Benefits Responsibilities- Perform in-depth analysis and optimization for current/next-generation NVIDIA GPUs
- Create and optimize core parallel algorithms and reference codes for NVIDIA GPUs
- Analyze hardware-software interactions on core algorithms, programming models, and applications
- Collaborate with hardware design, software engineering, product, and research teams to guide direction of accelerated computing
- Dive into accelerated computing applications to enable software-hardware co-design
- Document and present work through white papers, conference publications, blog posts, and patent applications as appropriate
Key requirements- MS or Ph.D. in Computer Science, Computer Engineering or Electrical Engineering, or equivalent experience
- 6+ years of relevant work experience
- Strong mathematical fundamentals (linear algebra, numerical methods)
- Passion for performance optimization
- Hands-on experience with massively parallel GPU programming model (CUDA or OpenCL)
- Familiarity with multi-node communication APIs (MPI or OpenSHMEM/NVSHMEM) is a plus
- Strong knowledge of C and C++ with good software design skills; familiarity with threading APIs and Unix IPC is a plus
- Familiarity with Python is a plus
- Good communication, organization, problem-solving, time management, and prioritization skills
- communication
- organization
- problem solving
- CUDA
- OpenCL
- MPI