Overview In this role you will bridge customers, engineering, and AI inference services to drive production outcomes on Tenstorrent AI systems. You'll deploy and operate production code with high autonomy, explaining trade-offs to both customer leadership and engineering teams. You'll work directly with customers to diagnose challenges across the inference stack and contribute measurable improvements. This is a hands-on, impact-driven engineering position at a remote-first company with hubs in North America.
Compensation / Benefits- competitive compensation package
- remote work flexibility (North America)
- equal opportunity employer
- global hub locations
Responsibilities- Contribute production code and operate deployments for AI inference workloads
- Serve as the continuity link between customers, engineering, and AI inference products
- Explain trade-offs and present recommendations to both customer leaders and engineering teams
- Debug across the full inference stack from requests to serving layer and kernel dispatch when needed
- Provide feedback via pull requests, reproducible code, benchmarks, and telemetry data
- Collaborate with customers to understand challenges and craft effective solutions
- Scale disaggregated inference services on Kubernetes while balancing performance and reliability
Key requirements- 5+ years of relevant software engineering experience (e.g., Applied Engineer, ML/AI/Platform/Infrastructure/SRE roles)
- Kubernetes and Helm experience at multi-node, HPC, or AI cluster scale
- Experience with observability and automation (Prometheus, Grafana, OpenTelemetry)
- Experience with LLM inference serving engines and technologies (e.g., vLLM, SGLang, Mooncake, NIM, Dynamo, LMCache)
- Customer-focused communication
- Collaborative problem-solving
- Ability to translate ambiguous requirements into acceptance criteria
- Kubernetes and Helm at scale
- Observability tooling (Prometheus, Grafana, OpenTelemetry)
- LLM inference serving technologies (vLLM, SGLang, Mooncake, NIM, Dynamo, LMCache)