×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff, Post-Training

Job in Menlo Park, San Mateo County, California, 94029, USA
Listing for: Goaly
Full Time, Apprenticeship/Internship position
Listed on 2026-09-03
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 230000 USD Yearly USD 150000.00 230000.00 YEAR
Job Description & How to Apply Below

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.

About the role

You will own the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning and related post-training experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role. The best candidates are equally comfortable forming hypotheses, writing high-quality code, operating training pipelines, and investigating why a model or metric moved. You will work closely with RL systems, training, inference, product, and domain experts; when infrastructure slows the science, you will help improve the infrastructure rather than treating it as someone else's problem.

What you'll do
  • Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.

  • Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.

  • Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.

  • Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect aggregate metrics to concrete model failures.

  • Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.

  • Analyze trajectories and model behavior to find reward hacking, shortcut learning, mode collapse, distribution gaps, and other failure modes; turn those findings into targeted experiments.

  • Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.

  • Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training, and data movement.

  • Translate successful research ideas into stable, repeatable pipelines and help set the team's longer-term post-training roadmap.

You may be a good fit if you have
  • Strong Python and software-engineering skills, including the ability to turn ambiguous research ideas into reliable experimental systems.

  • Hands-on experience training, fine-tuning, or evaluating modern language models, or an exceptional record in a closely related ML research area.

  • Solid understanding of deep learning and optimization, plus enough reinforcement-learning intuition to reason about policies, rewards, sampling, credit assignment, and evaluation bias.

  • Excellent experimental judgment: you define controls, inspect data, validate metrics, keep results reproducible, and distinguish a real improvement from noise or leakage.

  • Ability to debug across model behavior, data, code, and distributed infrastructure without losing sight of the user-facing capability being improved.

  • Clear written and verbal communication and a track record of productive collaboration across research and engineering.

Strong pluses
  • Experience with RLHF, reinforcement fine-tuning, preference optimization, reward or verifier modeling, or large-scale online sampling.

  • Experience building agent environments, secure sandboxes, coding benchmarks, tool-use tasks, or long-horizon evaluations.

  • Familiarity with PyTorch or JAX and distributed ML systems; experience with frameworks such as FSDP, Megatron, Deep Speed, Ray, veRL, or related stacks.

  • A record of influential research, open-source contributions, technically ambitious independent projects, or production model launches.

How we work
  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location,…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary