Overview In this role you will develop high-fidelity performance models to guide the Nemotron family of models and their deployment on target platforms. You will drive co-design decisions across AI research, framework, compiler, and hardware teams to achieve Pareto-optimal trade-offs between accuracy, throughput, and interactivity. You will translate modeling insights into architectural choices that scale in production and influence future software and hardware roadmaps. This is a chance to shape the next generation of GenAI systems at scale, with a strong emphasis on measurable impact and cross-functional collaboration.
Compensation / Benefits Responsibilities- Develop high-fidelity analytical performance models to prototype emerging algorithmic techniques and hardware optimizations for Nemotron models
- Prioritize features to guide future software and hardware roadmap based on detailed performance modeling and analysis
- Model end-to-end performance impact of GenAI workflows (e.g., Speculative Decoding, Agentic Pipelines, Inference-time compute scaling, RL) to forecast datacenter needs
- Collaborate with DL researchers, hardware architects, and software engineers to stay aligned with the latest DL research and cross-functional goals
Key requirements- Master's degree in Computer Science, Electrical Engineering or related fields
- Strong background in computer architecture, roofline modeling, queuing theory and statistical performance analysis techniques
- Solid understanding of ML fundamentals, model parallelism and inference serving techniques
- Proficiency in Python (and optionally C++) for simulator design and data analysis
- 3+ years of hands-on experience in system evaluation of AI/ML workloads, modeling and optimizations for AI
- Experience with deep learning frameworks like PyTorch, TRT-LLM, VLLM, SGLang
- Growth mindset and pragmatic "measure, iterate, deliver" approach
- Growth mindset
- Ability to distill complex analyses for both technical and non-technical audiences
- Collaborative by nature with cross-functional teams
- Python
- C++ (optional)
- PyTorch