Overview As a Senior Applied Research Scientist on NVIDIA's Curator team, you will advance deep learning models and data pipelines for multi-modal content extraction at petabyte scale. You will collaborate with researchers, ML engineers, and MLOps to improve curation quality for foundation-model training, including Nemotron Parse capabilities. You'll influence data pipelines, evaluation methodologies, and production readiness, shaping how large models learn from diverse documents, images, audio, and video. This is a mission-driven role at scale, enabling cutting-edge open LLMs.
Compensation / Benefits- equity and benefits
- base salary + bonus/equity potential
- remote work options
- competitive salaries
- flexible location NA/EU time zones
- comprehensive benefits package
Responsibilities- Develop efficient multi-modal data extraction and curation models and pipelines
- Build petabyte-scale extraction and content deduplication workflows (document/html parsing, fuzzy/near-duplicate, semantic, and substring deduplication)
- Scale curation methods across hundred-node GPU clusters to improve training set quality
- Design datasets, metrics, experiments, and validation scripts to standardize research methodologies
- Assist ML engineers in deploying pipelines to production via NVIDIA Inference Microservices (NIMs) and blueprints
- Produce technical content (papers, blogs, docs, training materials) to educate customers
- Stay current with data curation advances in academia and industry
Key requirements- Master's in Data Curation or equivalent experience in data curation, document AI, information retrieval or multimodal research
- Hands-on experience with computer vision and document-extraction models (layout analysis, OCR, table/figure/formula extraction)
- 10+ years in developing multimodal systems across models and platforms
- Experience managing distributed data frameworks (Ray, Spark, Dask) and deploying large multi-node ML tasks in production
- Strong Python skills and hands-on PyTorch (or similar) experience
- Ability to communicate ideas clearly via blogs, papers, kernels, GitHub
- Excellent communication and collaboration in a distributed team; mentoring juniors a plus
- Kaggle Grandmaster status or strong competition results a plus
- Remote-friendly; NA/EU time zones preferred
- Knowledge of batching, streaming, and scaling ingestion pipelines
- excellent communication and interpersonal skills
- mentoring junior engineers and interns
- ability to work in a dynamic, user-focused distributed team
- PyTorch
- Python
- OCR