Overview In this role you will set the technical direction for synthetic data generation across NVIDIA's frontier model efforts. You'll build open-source libraries in the NeMo ecosystem to generate diverse synthetic datasets for text, code, structured, and multimodal data to support pre- and post-training of LLMs like Nemotron. You will lead hands-on software engineering and applied research, collaborating with research, engineering, product, and external labs. You'll scale data pipelines, advance multimodal synthesis, and champion privacy-preserving strategies while publishing impactful research and mentoring the team.
Compensation / Benefits- equity
- benefits
- remote-friendly options
- base salary with defined range
- opportunity for growth and impact
Responsibilities- Define and build open-source data-generation libraries for NeMo
- Build and scale LLM-based data generation pipelines with quality evaluation
- Pioneer synthetic trajectories and tool-use data for reinforcement learning and reward modeling
- Advance multimodal synthetic data (image, document, video, audio)
- Implement privacy-preserving synthesis (differential privacy, anonymization)
- Develop and maintain APIs, documentation, and SDKs
- Publish research at top conferences and mentor scientists and engineers
Key requirements- PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience
- 15+ years of engineering and research experience in synthetic data generation, generative modeling, multimodal ML
- Deep understanding of LLMs and data's role in pre-training, post-training, and RL
- Proven track record of developing or maintaining widely used software libraries
- Experience building and optimizing scalable data pipelines for large-scale model training
- Strong publication record at NeurIPS, ICML, ICLR, ACL or equivalent
- collaboration across research, engineering, and product
- mentorship of scientists and engineers
- strong communication and documentation
- LLMs and data shaping for pre-training/post-training/RL
- synthetic data generation for text, code, structured, multimodal data
- multimodal data generation (vision-language, documents, video, audio)