Overview In this Principal Architect role, you will lead the research agenda for NVIDIA's AI system communications at scale, spanning GPUs, DPUs, NICs, and storage. You'll define long-term technical vision, translate research breakthroughs into production-grade software, and mentor senior engineers. The role blends systems software, high-performance networking, and hardware-software co-design to push data movement bottlenecks at the edge of AI infrastructure. You'll engage with industry forums and standards bodies and influence product strategy. This is a strategic, hands-on opportunity to shape how AI systems communicate across heterogeneous components.
Compensation / Benefits- equity
- comprehensive benefits package
- competitive salary
- opportunity to influence product roadmap
- work with leading AI infrastructure
- remote/hybrid options where applicable
Responsibilities- Set long-term technical vision for distributed AI communication systems (GPU-to-GPU, GPU-to-storage, cross-node data movement)
- Prototype and evaluate next-generation networking solutions over RDMA, NVLink, and GPUDirect
- Drive hardware-software co-optimization across GPU, DPU, NIC, and network switch
- Investigate bottlenecks in large-scale AI communication runtimes (KV cache transfer, disaggregated prefill/decode, model parallelism)
- Integrate networking capabilities into AI serving stacks (e.g., vLLM, SGLang, TensorRT-LLM)
- Publish findings, represent NVIDIA in industry forums and standards bodies, and mentor engineers
Key requirements- 15+ years in systems software and/or networking with deep expertise in high-performance networking (InfiniBand, RoCE, RDMA, NVLink)
- Proven track record of delivering cross-team technical initiatives from research concept to production
- MS, PhD or equivalent experience in Computer Science, Computer Engineering, Electrical Engineering, or related field
- Deep understanding of computer architecture, memory hierarchies, DMA engines, and OS-level networking
- Knowledge of ML systems concepts (transformers, KV cache, model parallelism, distributed training/inference)
- Proficiency in C, C++, Rust, and Python
- Strategic thinking and influence at senior levels
- Mentoring and leadership across teams
- Excellent collaboration and communication
- RDMA, NVLink, GPUDirect
- InfiniBand, RoCE, NVLink
- NIXL, NCCL, UCX, MPI, NVSHMEM