Overview Lead Machine Learning Engineer at Cisco AI Research, based in Seattle or San Jose (hybrid). You will design scalable data pipelines and ML systems powering LLMs and AI models, focusing on high quality training data at scale. You'll collaborate with researchers and engineers to determine data needs and measure impact on model performance. This role blends ML engineering with data engineering to drive end to end data-to-model workflows. You'll work on impactful data strategies, from labeling to synthetic data and evaluation, shaping production AI solutions.
Compensation / Benefits- medical, dental and vision insurance
- 401(k) with match
- paid parental leave
- paid time off including holidays and vacation
- sick time and flexible vacation policy
- volunteer days
Responsibilities- Design, build, and maintain scalable data pipelines for ML and LLM lifecycles from data ingestion to deployment
- Architect and manage human in the loop labeling workflows including quality control and feedback integration
- Develop scalable synthetic data generation, filtering, and validation to improve dataset quality
- Leverage LLMs and ML techniques to automate data generation, labeling, scoring, and evaluation
- Establish systems to measure and mitigate dataset failure modes (bias, contamination, distribution shifts) and tie dataset changes to model performance
- Collaborate with researchers and ML engineers to define dataset requirements for fine tuning, preference learning, and agent development
- Provide technical direction on infrastructure, compute, and storage; promote engineering excellence through design reviews and mentorship
Key requirements- Bachelor's degree in a STEM field with 8+ years of relevant experience, OR Master's in STEM with 6+ years, OR PhD with 3+ years of research experience
- 3+ years building, curating, and scaling ML datasets for training/evaluation
- 5+ years professional programming in Python, C++, or Go
- 5+ years experience with ML frameworks (PyTorch, TensorFlow) in production or research
- strong collaboration with researchers, engineers, and product stakeholders
- ability to communicate complex data concepts clearly
- mentorship and technical leadership
- Data pipeline design and data engineering for ML/LLMs
- Human in the loop labeling systems and quality control
- Synthetic data generation, augmentation, and validation