Research Scientist, Reinforcement Learning
Listed on 2026-08-01
-
Software Development
Data Scientist, Machine Learning/ ML Engineer
Member of Technical Staff, Reinforcement Learning Research
Our client is a well-funded, early-stage AI lab building a real-time, multimodal AI system designed to interact with the natural timing, emotion, and responsiveness of a human. The team is lean, highly credentialed, and moving fast on problems that are genuinely unsolved.
About the RoleOwn RL and post-training for large-scale multimodal models at a frontier lab, from building the stack at 0→1 to scaling it to production. This is broader than a traditional RL algorithms role - you'll develop post-training methods and build the infrastructure needed to run them. The work spans rollout generation, reward modeling, policy optimization, evaluation, data feedback loops, serving, observability, and distributed execution.
Responsibilities- Build the RL/post-training stack from scratch: rollout generation, policy optimization, reward and reference model serving, data feedback loops, evaluation, checkpointing, and observability.
- Develop and scale post-training methods including PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
- Design systems abstractions connecting research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
- Build evaluation and feedback loops for multimodal behavior: turn-taking, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
- Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, and research iteration speed.
- Evolve the platform as algorithms, model architectures, reward definitions, data sources, and evaluation methods change.
- PhD in ML, RL, CS, or a related field - completed or in final stretch
- Deep knowledge of RL and post-training methods: policy optimization, reward modeling, preference optimization, rejection sampling, KL control, and data feedback loops
- Ability to reason about training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch
- Hands-on exposure to RL/post-training pipelines through research, internships, or open-source work; familiarity with frameworks such as verl, OpenRLHF, ms-swift, or equivalent, and rollout serving systems such as vLLM or SGLang
- Strong software engineering fundamentals with the appetite to build real systems, not just prototypes
- Must be willing to work on-site 5 days/week in Seattle, WA (relocation support available)
- First-author publications at tier-1 ML journal/conference
- Internship experience at a top AI lab doing RL or post-training work
- Prior 0→1 experience building post-training systems, RL pipelines, agent training infrastructure, or evaluation platforms
- Hands-on experience with multimodal post-training for audio, video, or language models - particularly long-context or real-time interactive systems
- Experience with adjacent areas: distributed pretraining, inference serving, data infrastructure, simulation, or human/AI feedback collection
- Substantial open-source contributions in RL, post-training, alignment, or ML systems
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).