Overview In this role you will advance AI-powered drug discovery by developing scalable ML systems and improving large language models for biomolecular design. You'll bridge engineering and research, turning scientific ideas into production-ready tooling. You'll work with cross-functional scientists to translate biology into machine learning objectives, benchmarks, and data pipelines. This is an opportunity to shape how data and AI accelerate medicines for patients worldwide.
Compensation / Benefits- discretionary annual bonus
- benefits package per policy
- opportunity to work at biotech leader
- relocation not offered
- salary range varies by location
- full-time employment
Responsibilities- Design, implement, and scale large-scale distributed ML systems and core infrastructure
- Develop strategies to improve model performance on scientific tasks and long-horizon reasoning
- Translate biological/chemical knowledge into ML objectives, signals, and evaluation criteria
- Design evaluation methodologies and collaborate with domain experts to establish benchmarks and data quality
- Collaborate with researchers to translate ideas into scalable, production-ready systems
- Maintain training infrastructure and data pipelines to ensure reliable experiments on clusters
- Work with senior scientists to implement novel algorithms and prototype research into software
Key requirements- BS/MS in Computer Science, Statistics, Mathematics, Physics, or related quantitative field with 2+ years of experience, or PhD with 0-2 years experience
- Experience developing and training large-scale ML models including domain knowledge enhancement and alignment
- Strong software engineering skills and experience with high-performance computing systems
- Publication record in top-tier venues (e.g., NeurIPS, ICLR, ICML)
- Experience collaborating with researchers and translating research into production
- collaboration with cross-functional teams
- ability to translate domain knowledge into ML problems
- communication of complex ideas to diverse stakeholders
- LLM development and training
- distributed ML systems
- training infrastructure and data pipelines