Machine Learning Scientist
Spotter · Los Angeles, CA
📍 Culver City, California, United Statesvia greenhousePosted 2026-08-06
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to Spotter.
Overview
Spotter empowers the world's best Creators with capital, data, and insights to scale their programming into sustainable media businesses. Through these partnerships, Spotter helps brands partner with creator-led franchises to unlock growth, amplify impact, and build lasting cultural relevance.
Spotter has already deployed over $1 billion to YouTube Creators to reinvest in themselves and accelerate their growth. With a premium catalog that spans over 725,000 videos , Spotter generates more than 88 billion monthly watch-time minutes, delivering a unique scaled media solution to Advertisers and Ad Agencies that is transparent, efficient, and 100% brand safe. For more information about Spotter, please visit https://spotter.com .
Overview
We're looking for a talented and intensely curious Machine Learning Scientist with deep expertise in building and deploying production machine learning models, particularly reinforcement learning, contextual bandits, and adaptive learning systems, along with deep learning, ranking, personalization, and recommendation systems. You thrive in a fast-paced startup environment and are motivated by building models that don't just perform well in experiments, they ship to production and create real value for YouTube Creators.
In this role, you'll train, evaluate, optimize, and deploy a wide range of machine learning models, from contextual bandits and sequential decision-making systems to neural networks, ranking systems, recommendation models, and traditional machine learning approaches. You're passionate about staying at the forefront of AI and machine learning, especially in areas where models learn from feedback, adapt over time, and improve real-world product outcomes.
We're a team of builders who value continuous learning, rapid experimentation, and delivering AI solutions that make a measurable difference for Creators. If you enjoy solving complex problems, iterating quickly, and building intelligent products that help the world's top YouTube Creators work smarter and create better content, you'll thrive at Spotter.
What You’ll Do
You'll develop machine learning models that move beyond experimentation and into production, where they directly improve Creator workflows and product experiences. Working alongside Analytics, Product, and Engineering, you'll help develop intelligent systems that improve how Creators discover insights, make decisions, and create content.
Your work may include:
Designing, training, evaluating, optimizing, and deploying production reinforcement learning, contextual bandit, and online learning systems that improve product outcomes.
Creating systems that balance exploration and exploitation, short-term performance and long-term value, and multiple competing product objectives.
Developing reward models, feedback models, and objective functions that translate noisy, sparse, delayed, or implicit signals into reliable model training and evaluation targets, and diagnosing and mitigating reward hacking and feedback loops in deployed systems.
Applying offline policy evaluation and counterfactual techniques, such as inverse propensity scoring, doubly robust estimation, and replay evaluation, to reason about model changes before and after deployment.
Working with logged interaction data to understand user behavior, evaluate model performance, improve decision quality, and reduce bias in model evaluation.
Designing experiments to evaluate model performance, measure product impact, and continuously improve production systems.
Building scalable model training, evaluation, deployment, and inference pipelines.
Optimizing models for accuracy, latency, scalability, reliability, and production maintainability.
Working with structured and unstructured datasets using Python and SQL.
Collaborating closely with Product and Engineering to translate customer problems into machine learning solutions.
Staying current with advances in reinforcement learning, bandits, recommendation systems, ranking, personalization, deep learning, experimentation, and production ML, and thoughtfully applying new techniques where they create measurable value.
Who You Are
Required Skills & Experience
Master's degree or PhD in Computer Science, Statistics, Applied Mathematics, Electrical Engineering, Physics, or another quantitative field.
5+ years building, evaluating, and deploying machine learning models in production environments.
Experience with reinforcement learning or contextual bandit systems gained through graduate coursework, academic research, or hands-on industry experience. Candidates with experience building and deploying these systems in production, from problem formulation through offline evaluation to live deployment, are strongly preferred.
Solid grasp of core RL training objectives and loss functions, including temporal-difference and Bellman error losses (Q-learning, DQN), policy gradient objectives (REINFORCE, actor-critic advantage estimation), and clipped surrogate objectives (PPO, TRPO), with an understanding of when each applies and how they behave in training.
Practical experience with bandit and reinforcement learning methods such as Thompson sampling, UCB or LinUCB, neural bandits, non-stationary bandits, policy gradients, actor-critic methods, or Q-learning.
Ability to design reward functions and objective trade-offs for systems optimizing long-horizon outcomes, including diagnosing and mitigating reward hacking and feedback loops.
Knowledge of off-policy and counterfactual evaluation, such as inverse propensity scoring (IPS), self-normalized IPS, doubly robust estimators, and replay evaluation, and with counterfactual learning from logged bandit feedback, including propensity logging.
Experience working with logged interaction data, behavioral data, or feedback signals to train, evaluate, and improve models.
Track record of designing experiments and using data to improve mod
More Los Angeles, CA jobs
Los Angeles, CA jobs · Browse all locations