×
Register Here to Apply for Jobs or Post Jobs. X

Applied Scientist, Reinforcement Learning; Mid, Senior

Job in Menlo Park, San Mateo County, California, 94025, USA
Listing for: Hippocratic AI
Full Time position
Listed on 2026-08-25
Job specializations:
  • IT/Tech
    Machine Learning/ ML Engineer, AI Evaluation
Job Description & How to Apply Below
Position: Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Role Mission

As HAI's LLM Post-Training Applied Scientist, you will own the reinforcement learning and on-policy distillation pipeline that transforms raw model capability into reliable, safe clinical behavior. Your post-training methods will directly determine how our AI agents reason through complex clinical scenarios, handle safety-critical decisions, and ultimately impact millions of patient interactions. This role exists because post-training is where capability becomes trustworthiness—and in healthcare, that's everything.

What

You Will Accomplish

Own your first major outcome:
By day 90, you will have shipped a post-training improvement that meaningfully advances model performance on a critical clinical capability (clinical reasoning, safety alignment, or task completion), evaluated the gains rigorously, and contributed that learning to our post-training roadmap.

Drive lasting impact:
At 12 months, you will have designed and shipped multiple post-training methods that measurably improve our models' clinical safety and reasoning, built reusable infrastructure (reward models, verifiers, evaluation frameworks) that accelerate future post-training work, published your research or contributed to HAI's intellectual property, and directly shaped how our deployed models behave in production healthcare environments.

The Team

You'll work alongside ML researchers, engineers, clinicians, and safety experts who are obsessed with building trustworthy AI. This is a highly technical team that values rigor, collaboration across disciplines, and solving the hardest problems in AI safety and alignment. You'll have direct influence on model architecture and training decisions.

What You'll Do
  • Design and implement RL and OPD post-training methods including RLHF, RLVR, on-policy distillation, and novel approaches tailored to healthcare AI—selecting the right methods for different clinical reasoning and safety challenges

  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines that provide reliable training signals for post-training, ensuring they capture what truly matters in clinical contexts (accuracy, safety, patient experience)

  • Develop conversational AI environments and simulations for healthcare RL training—creating synthetic clinical scenarios and datasets that enable safe, scalable post-training without relying solely on human feedback

  • Automate post-training research loops using agents and tooling to systematically explore hyperparameters, methods, and data strategies—turning post-training into a scientific, reproducible process

  • Run rigorous experiments and analysis to understand what drives post-training gains, isolate the contributions of different components, and build intuition about what works in healthcare contexts

  • Collaborate with research, engineering, and clinical teams to translate clinical requirements into post-training objectives, validate improvements against real-world metrics, and scale successful methods to production

Location Requirement

We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Palo Alto office five days a week, unless otherwise specified.

Applied Scientist, Reinforcement Learning

Required Qualifications

  • Master's degree in Computer Science, Machine Learning, or a related field

  • 5+ years of professional experience in NLP, LLM training, or reinforcement learning

  • 2+ years of hands-on experience with RL for LLM post-training

  • Proficiency in Python and PyTorch for large-scale training

  • Demonstrated experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods

  • Experience training or fine-tuning models at scale (50B+ parameters)

Preferred Qualifications

  • Publications at top-tier ML venues (NeurIPS, ICML, ICLR, ACL, EMNLP)

  • Healthcare or regulated domain experience

  • Experience with distributed training frameworks (FSDP, Deep Speed, vLLM)

  • Familiarity with safety alignment and interpretability research

Our comprehensive compensation package is designed to reward your expertise and includes both a competitive base salary and valuable stock options. Individual offers are determined based on a variety of…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary