Overview In this role you apply reinforcement learning and related decision-making methods to optimize real-time advertising decisions across targeting, bidding, and personalization. You will translate research into production-ready models that operate at high throughput and low latency, and you'll work with cross-functional partners to measure impact through offline and online experiments. Your work shapes how Viant's autonomous advertising systems improve campaign performance and auction efficiency. You'll contribute to a rigorous, research-oriented culture while delivering practical, measurable outcomes.
Compensation / Benefits- fully paid health insurance
- paid parental leave
- unlimited PTO
Responsibilities- Develop, train, and evaluate RL, contextual bandit, ranking, and prediction models for ad optimization, bid optimization, targeting, and personalization
- Study auction dynamics, exploration/exploitation, budgeting, pacing, and reward design to improve real-time advertising decisions
- Translate research ideas into production-ready models operating at high throughput and low latency
- Design and analyze offline and online experiments, including counterfactual and off-policy evaluation
- Collaborate with engineers to deploy, monitor, retrain, and improve models in production while addressing data leakage, drift, and market changes
- Apply quantitative reasoning and statistical modeling to CTR, conversion, ROAS, targeting, attribution, identity, and measurement problems
- Collaborate with scientists, engineers, and product partners to define objectives, labels, loss functions, reward signals, evaluation metrics, and delivery plans
- Contribute to a rigorous, research-oriented team culture through technical communication, code and model reviews, experimentation, and knowledge sharing
Key requirements- 1-3 years of experience developing and applying machine learning models in production or research with measurable outcomes
- Strong foundation in ML, deep learning, probability, statistics, and optimization; practical experience with Python and PyTorch or TensorFlow
- Experience with reinforcement learning, contextual bandits, sequential decision-making, recommender systems, online experimentation or related methods
- Ability to formulate ML problems precisely (objectives, labels, features, loss/reward, evaluation, experimental design)
- Experience analyzing large-scale data and communicating findings to scientists, engineers, and cross-functional partners
- Interest in moving beyond offline accuracy to improve real-world production decisions
- Experience with RL in advertising, marketplaces, recommender systems, robotics, games, or other sequential decision-making environments
- Exposure to contextual bandits, off-policy/counterfactual evaluation, causal inference, auction theory, or online experimentation
- clear communication of technical findings to cross-functional partners
- collaboration with scientists, engineers, and product teams
- problem solving and quantitative reasoning
- reinforcement learning
- contextual bandits
- online experimentation