×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Research Engineer - RL and Post Training

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: BioStack Platforms
Full Time, Apprenticeship/Internship position
Listed on 2026-08-22
Job specializations:
  • Business
    AI Evaluation
Salary/Wage Range or Industry Benchmark: 200000 - 350000 USD Yearly USD 200000.00 350000.00 YEAR
Job Description & How to Apply Below

Bio Stack is building the data layer for AI-native healthcare and drug discovery. We work with leading AI labs, human data companies, and frontier biotech teams to source, structure, and deliver high-value clinical and preclinical datasets for model training, evaluation, and deployment.

We sit at the intersection of healthcare, frontier AI, and data infrastructure. Our work spans medical institutions, clinics, imaging centers, and data partners globally, turning messy real-world clinical workflows into AI-ready products that matter.

The long‑term vision is to make high-quality healthcare accessible to everyone

and radically improve drug discovery by linking real-world healthcare data with genomics,

imaging, biomarkers, and experimental data. This creates a foundation for AI systems that can

learn from millions of patient journeys, understand why treatments work for some patients and

fail for others, personalize care based on clinical and genomic context, identify the right

interventions earlier, and uncover new therapeutic opportunities from the connection between

biology and real-world outcomes.

Bio Stack is backed by PeakXV, Y Combinator, Afore Capital, SV Angel as well as high-profile angels from OpenAI, Meta and Google Deep Mind.

About the Role

As an RL Engineer at Bio Stack, you will build reinforcement learning environments and post‑training systems for healthcare AI.

Bio Stack is building the data and environment layer for medical AI: sourcing high‑value clinical data, turning it into model‑ready workflows, and building tasks, rewards, verifiers, benchmarks, and agent environments where models can learn against meaningful and measurable outcomes.

You will work across the full RL loop — from environment and reward design to training, evaluation, and iteration. Projects may span clinical reasoning, longitudinal patient care, diagnostic decision‑making, chronic disease management, and biomedical research.

This is a hands‑on engineering role. You will build environments, run experiments, train models and agents, analyze failures, and improve the data and feedback signals that determine what models learn.

Strong judgment around data is particularly important. You should be able to determine whether a dataset has the signal quality, label fidelity, coverage, diversity, and clinical relevance required to support useful training tasks, rewards, and evaluations.

Prior healthcare experience is not required.

This is a full‑time in‑person role based in San Francisco, CA.

What you will do:
  • Build healthcare‑specific RL environments, including tasks, action spaces/tool interfaces, reward functions, verifiers, and evaluation harnesses.
  • Run post‑training experiments on language models and agents using techniques such as SFT, RLVR, RLHF/RLAIF, and reward modeling.
  • Turn clinical and biomedical datasets into training environments with measurable, verifiable outcomes.
  • Design rewards and verifiers that capture correctness across clinical reasoning and longitudinal decision‑making tasks.
  • Train and evaluate multi‑step agents operating across patient histories, clinical tools, and structured/unstructured medical data.
  • Build scalable pipelines for rollouts, training, evaluation, experiment tracking, and dataset iteration.
  • Analyze model failures and use them to improve environments, rewards, datasets, and subsequent training runs.
You might thrive in this role if:
  • You are excited by the idea of applying frontier RL methods to healthcare, medicine, and biological data.
  • You have experience with reinforcement learning, language model post‑training, agent environments, reward modeling, evaluation, or related ML systems.
  • You have strong judgment around data and can assess whether a dataset has sufficient signal quality, label fidelity, coverage, longitudinal depth, and clinical relevance to support meaningful training tasks, environments, rewards, and evaluations.
  • You can move quickly from research concept to working prototype, then iterate based on empirical results.
  • You are comfortable designing controlled experiments, building baselines, and drawing trustworthy conclusions from noisy real‑world data.
  • You are comfortable working in large ML…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary