Senior Applied Scientist, Reinforcement Learning
Listed on 2026-07-31
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Data Scientist
About the Role
LLM post-training is where raw capability becomes reliable, safe behavior — and in healthcare, the stakes are as high as they get. You'll own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline end to end, to improve our models' clinical reasoning, safety, and alignment. Your models will be deployed to interact with millions of patients across diverse clinical use cases.
WhatYou'll Do
Design RL and OPD post-training methods (RLHF, RLVR, OPD, etc.)
Build and evaluate reward models, verifiers, and LLM-as-judge pipelines
Develop conversational AI environments and simulations for healthcare RL training with synthetic data
Automate post-training loops with agents (auto-research)
Run rigorous experiments to understand what drives post-training gains
Collaborate with research, engineering, and clinical teams
MS or PhD in CS or relevant field
3+ years or experience in NLP, LLM training, or RL
1+ years experience in RL for LLM post-training
Experience with large-scale (50B+ parameter and multi-node) LLM training
Strong Python and PyTorch coding skills
Experience with RLHF, RLVR, LLM-as-judge or similar methods for LLM post-training
Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @ email addresses. We will never request payment or sensitive personal information during the hiring process.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).