Overview In this role you will lead the architectural definition and improvement of CPU cache hierarchies and interconnects. You will create and maintain high-fidelity, cycle-accurate performance models that guide data movement across silicon, supporting automotive and data center needs. You will work with cross-functional teams to drive architectural trade-offs and ensure models align with silicon and software performance. This is a chance to shape scalable, high-performance systems at NVIDIA, balancing automotive reliability with data-center efficiency.
Compensation / Benefits Responsibilities- Develop and maintain cycle-accurate performance models (C++/SystemC) for coherent interconnects and large-scale shared caches
- Model and analyze performance bottlenecks across scales from automotive SoCs to data center architectures
- Evaluate performance impact of coherency protocols (CHI, ACE, or proprietary) and snooping filters
- Run and analyze industry benchmarks (SPEC, MLPerf, Automotive suites) to drive architectural trade-offs
- Collaborate with build/verification to correlate models with silicon; work with software to optimize drivers for hardware topology
Key requirements- Master's in Computer Engineering, Electrical Engineering, or Computer Science (or equivalent experience) with 5+ years of experience
- Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols
- Proven experience in C++ or SystemC for cycle-accurate or functional modeling
- Proficiency in Python or similar scripting languages for data processing, visualizations, and automating sweeps
- Understanding of NoC topologies (Mesh, Ring, Torus), credit-based flow control, and arbitration logic
- systems-thinking
- cross-domain collaboration
- analytical mindset
- C++
- SystemC
- Python