Overview As an Architecture Leader, you shape the long-term roadmap for AI communication libraries across NVIDIA's platforms. You drive scalable architectures for models spanning hundreds of thousands of nodes and partner with application teams to optimize primitives for AI and HPC workloads. You influence hardware-software co-design with silicon teams to meet trillion-parameter AI demands. This role offers the chance to impact large-scale systems, accelerate distributed AI, and work at the forefront of GPU-enabled computing.
Compensation / Benefits- equity
- comprehensive benefits
- competitive base salary
- remote options
- career growth opportunities
- exclusive engineering teams
Responsibilities- Define long-term roadmaps for communication libraries across NVIDIA platforms
- Lead development of next-gen communication primitives and collective algorithms
- Collaborate with application teams to design specialized primitives for AI and HPC libraries (NCCL, NIXL, NVSHMEM, UCC, UCX)
- Influence hardware specifications through hardware/software co-design for next-gen networking
- Develop high-fidelity models and simulators to predict system behavior under new workloads
Key requirements- Ph.D. or M.S. in Computer Science, Electrical Engineering, or related field with 12+ years in HPC or distributed deep learning
- Deep understanding of 3D parallelism (Data, Tensor, Pipeline) and ZeRO variants
- Proficiency with NCCL, UCX, UCC, NVSHMEM, or MPI; experience with RDMA, RoCE, and InfiniBand verbs
- In-depth knowledge of high-throughput inference engines and schedulers (TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo)
- Expertise in NVIDIA GPU memory hierarchy (HBM3e/HBM4, L2 cache) and CUDA programming models
- collaboration across cross-functional teams
- strong problem-solving and analytical thinking
- leadership and technical mentorship
- NCCL
- UCX
- UCC