Founding AI Researcher, RL
Listed on 2026-10-02
-
IT/Tech
AI Evaluation, Machine Learning/ ML Engineer
We’re building toward a world where every company can become its own AI lab.
Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.
Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.
About the roleYou will own the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning and related post-training experiments, diagnose model behavior, and convert results into better recipes and production models.
This is a research-engineering role. The best candidates are equally comfortable forming hypotheses, writing high-quality code, operating training pipelines, and investigating why a model or metric moved. You will work closely with RL systems, training, inference, product, and domain experts; when infrastructure slows the science, you will help improve the infrastructure rather than treating it as someone else’s problem.
What you'll doDesign and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.
Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.
Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.
Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect aggregate metrics to concrete model failures.
Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.
Analyze trajectories and model behavior to find reward hacking, shortcut learning, mode collapse, distribution gaps, and other failure modes; turn those findings into targeted experiments.
Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.
Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training, and data movement.
Translate successful research ideas into stable, repeatable pipelines and help set the team's longer-term post-training roadmap.
Strong Python and software-engineering skills, including the ability to turn ambiguous research ideas into reliable experimental systems.
Hands‑on experience training, fine‑tuning, or evaluating modern language models, or an exceptional record in a closely related ML research area.
Solid understanding of deep learning and optimization, plus enough reinforcement-learning intuition to reason about policies, rewards, sampling, credit assignment, and evaluation bias.
Excellent experimental judgment: you define controls, inspect data, validate metrics, keep results reproducible, and distinguish a real improvement from noise or leakage.
Ability to debug across model behavior, data, code, and distributed infrastructure without losing sight of the user-facing capability being improved.
Clear written and verbal communication and a…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).