Overview In this role you will define server architectures for AI training at cloud scale and translate workloads into detailed hardware specs. You'll lead ODM/JDM partners, drive validation from PCBA bring-up through rack integration, and own fleet quality metrics post-launch. Expect cross-functional collaboration across hardware, firmware, software, and operations to deliver scalable, debuggable, and serviceable AI infrastructure. Your work directly shapes the hardware that powers frontier models and large-scale inference.
Compensation / Benefits- Health insurance (medical, dental, vision)
- 401(k) matching
- RSUs and sign-on incentives
- Paid time off
- Parental leave
- EAP and mental health support
Responsibilities- Define server architectures based on workload demand and translate into designs and component specs
- Collaborate with interdisciplinary teams to deliver cohesive designs
- Lead ODM/JDM design reviews covering schematic, layout, BOM, and DFx
- Define and execute validation strategies from PCBA bring-up to rack integration (power sequencing, thermal, SI, accelerator interconnect)
- Own hardware debug during EVT/DVT/PVT, correlating failures across PCIe, power, memory, and GPUs
- Triage hardware issues at ODM facilities and datacenters; perform root cause analysis and corrective actions
- Own fleet quality metrics post-launch (annualized server/component failure rates) and drive design/process improvements
- Monitor telemetry and partner with test/automation teams to improve yield and reduce test dwell
- Collaborate with EC2 architecture teams to align on requirements and platform trade-offs
- Guide ODM/JDM partners through development milestones and production ramp
- Mentor junior engineers and contribute to hiring and knowledge sharing
- Travel up to 10% regionally/internationally to partner sites as needed
Key requirements- Bachelor's degree in Electrical Engineering or Computer Engineering, or equivalent
- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms
- Experience in server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems
- Experience leading hardware development through full product lifecycle (concept through production ramp)
- Data-driven decision making
- Collaborative mindset and cross-functional communication
- Mentorship and team development
- Thermal and mechanical design for high-density compute
- Power delivery and signaling integrity for GPU/accelerator platforms
- High-speed interconnects and PCIe/NVMe analysis