Overview As the Technical Lead for Compute Fleet Management, you will shape how Databricks consumes and optimizes compute across AWS, Azure, and GCP. You drive fleet provisioning at scale, ensure high availability and isolation, and own the architecture for a resilient, low-dependency compute platform. You work across teams to deliver impactful, cross-cloud initiatives that improve margin and customer experience. This role offers meaningful influence over scalable infrastructure used by thousands of customers.
Responsibilities- Pioneer fleet optimization by provisioning and pooling massive cloud resources for peak workload performance and isolation
- Deliver hyper-scale resilience with horizontal scaling and cloud-account failure tolerance
- Own the critical path by leading development of low-dependency systems to bootstrap and manage the compute platform
- Drive high availability (99.99%), and maximize utilization (60%+) while balancing cloud failures and performance
- Architect and enforce strong security and performance isolation across diverse workloads
- Lead cross-team, cross-layer strategic engineering initiatives from concept to execution
- Support GPU scaling for AI/ML workloads and multi-cloud distributed system operation across major clouds
Key requirements- Seasoned Principal Engineer with hands-on experience building and operating large-scale, mission-critical infrastructure in production
- Proven track record leading transformative, cross-team engineering initiatives
- Distributed systems mastery with experience on at least one major public cloud
- Able to influence, drive consensus, and lead large technical efforts across organizational boundaries
- Strong execution discipline in planning, tracking, and managing complex dependencies
- influence without authority
- cross-functional collaboration
- strategic thinking
- large-scale distributed systems
- multi-cloud (AWS, Azure, GCP) operations
- systems architecture for compute fleets