Overview In this role you will be a technical multiplier for AI labs and AI-native companies, bridging bleeding-edge infrastructure goals with Datadog's roadmap. You will lead high-level AI/LLM architecture discussions and influence stakeholders to realize scalable observability for training and deploying models. You'll collaborate across PSA, Sales, and Marketing to keep Datadog at the forefront of the AI stack and shape product directions. This role offers impact in shaping AI-native observability at scale and engaging with senior technical leaders.
Compensation / Benefits- best-in-breed onboarding
- generous global benefits
- mentor and buddy program
- stock equity and ESPP
- continuous professional development
- inclusive culture and community programs
Responsibilities- Demonstrate thought leadership in AI/LLM space and connect capabilities to business impact
- Advise founders, heads of infrastructure, and researchers on best practices and trends in AI/LLM
- Lead architecture reviews and design engagements for high-throughput AI workloads
- Identify AI-native technology shifts and feed insights back to Product Management; co-create observability integrations
- Collaborate with PSA, Sales, Sales Engineering, and Marketing to provide technical resources
- Assist leadership in recruiting and hiring for PSA and Field CTO teams
Key requirements- 10+ years of experience with at-scale distributed systems architecture, HPC, or large-scale infrastructure
- Deep familiarity with AI/LLM ecosystem, accelerator hardware (GPUs/TPUs), and modern orchestration frameworks
- Strategic thinking and ability to address wide technical and business challenges
- Proven experience influencing elite individual contributors, researchers, and technical founders
- Hands-on technologist with ability to white-board architectural solutions with senior engineers
- Strong presentation skills for large audiences and executives
- Excellent verbal and written communication linking product functionality to business value
- Willingness to travel up to 50%
- strategic thinking
- stakeholder management
- presentation and communication
- AI/LLM systems and observability (LLMO)
- high-performance and scalable infrastructure
- GPU/TPU acceleration