Overview In this role you advance safety post-training methods and interpretability to make frontier AI systems safer and more understandable for researchers and policymakers. You will work within Scale's policy research team to connect ML findings with safety standards and real world evaluation benchmarks. You'll design pipelines to study how training choices shape safety and alignment, and translate insights into practical mitigations. This position offers a mission driven opportunity to influence AI risk governance across industry, government, and academia.
Compensation / Benefits- comprehensive health, dental and vision coverage
- retirement benefits
- learning and development stipend
- generous PTO
- commuter stipend (possible)
Responsibilities- Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties
- Develop interpretability informed evaluations to reveal unsafe or undesirable model behaviors and guide mitigations
- Collaborate with policymakers, engineers, and researchers to translate findings into actionable safety standards, benchmarks, and best practices
Key requirements- Commitment to safe, trustworthy AI deployments
- Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches
- A track record of published research in machine learning, particularly in generative AI
- At least three years of experience addressing sophisticated ML problems in research or product development
- Strong written and verbal communication skills to operate in a cross-functional team
- Cross-functional collaboration
- Clear and effective communication
- Analytical thinking and problem solving
- Post-training methods
- Reinforcement learning techniques (RLHF, DPO, GRPO)
- Mechanistic interpretability (nice to have)