Overview In this role you will accelerate drug discovery by building and applying predictive models to transform complex biological and chemical data into actionable insights. You will operate at the intersection of data science, chemistry, and biology to support target discovery, compound optimization, and translational research. You'll contribute to a data-driven discovery ecosystem by delivering models, analyses, and reusable workflows in collaboration with scientists across disciplines. This position offers the opportunity to impact drug discovery at scale through rigorous ML methods and cross-functional partnerships.
Compensation / Benefits- competitive cash compensation
- robust equity awards
- strong benefits
- learning and development opportunities
- hybrid work model ()
- base pay disclosed for location
Responsibilities- Develop, implement and evaluate machine-learning models to support drug discovery questions (activity, selectivity, developability, target engagement, phenotypic outcomes)
- Perform exploratory data analysis and quality assessment on chemical, biological, imaging, and phenotypic datasets
- Prepare and integrate heterogeneous datasets (chemical structures, screening data, structural biology outputs, molecular simulations, high-content imaging)
- Apply supervised learning, deep learning, graph-based methods and ensemble approaches under guidance of leads
- Ensure robust validation, applicability, and limitations of models
- Collaborate with data engineering and ML engineering partners to enable reproducible workflows and integration into discovery pipelines
- Partner with medicinal chemists, biologists, and other researchers to translate questions into analyses and communicate results clearly
- Document methods, code, results, and assumptions to support reproducibility
Key requirements- Ph.D. in a related quantitative field or M.S. with relevant industry experience
- Typically 2-5 years of experience applying ML/data science to scientific datasets
- Experience developing, validating, and evaluating predictive or classification models
- Strong Python programming skills with NumPy, Pandas, SciPy
- Hands-on experience with ML frameworks (PyTorch, TensorFlow, scikit-learn)
- Experience with data visualization and working with noisy or incomplete data
- Ability to communicate technical work clearly and collaborate with cross-functional teams
- collaborative mindset
- clear communication
- curiosity about drug discovery
- Python programming
- NumPy
- Pandas