Member of Technical Staff, Research — Early Career(PHD
Listed on 2026-08-19
-
Research/Development
AI Evaluation, Research Scientist, AI Business & Operations
About us
We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.
About the roleYou will own research at the intersection of agentic reinforcement learning, post-training, evaluation, environments, and scaling. You will identify high-leverage questions, design and run decisive experiments, and translate results into model improvements and reusable systems.
This role is designed for researchers completing or recently completing a PhD, as well as candidates with an equivalent record of original research. It is not a purely academic position: strong candidates write excellent code, work closely with systems engineers, and care whether an idea survives realistic evaluation and production constraints.
What you’ll doFormulate high-leverage research questions about agentic capability and reliability, RL algorithms, reward and verifier design, exploration, curricula, environment design, task distributions, and scaling behavior.
Design rigorous experiments, ablations, controls, and evaluations that separate real model improvement from noise, data leakage, reward hacking, or benchmark overfitting.
Implement new methods in modern deep-learning frameworks and integrate them with production training, rollout, environment, and evaluation systems.
Build or improve datasets, agent environments, verifiers, and evaluations for domains such as coding, tool use, reasoning, long-horizon tasks, or computer interaction.
Analyze trajectories and model behavior, develop useful failure taxonomies, and turn observations into testable hypotheses and prioritized experiments.
Partner closely with Post-Training, RL Systems, Training, and Backend & Product engineers to scale promising ideas and expose them to realistic product constraints.
Communicate findings in clear internal documents and technical reviews; contribute to papers, technical reports, blog posts, or open-source releases when aligned with company goals.
Help shape the research roadmap by identifying compounding capabilities, reusable evaluation assets, and experiments that retire the most important uncertainties.
Completing or recently completed a PhD in computer science, machine learning, statistics, mathematics, or a related field—or an equivalent record of original, technically rigorous research.
A strong research record in machine learning, reinforcement learning, large language models, agents, or ML systems, demonstrated through publications, preprints, open-source work, or substantial independent projects.
Excellent Python skills and hands-on experience with a modern deep-learning framework such as PyTorch or JAX.
Experimental rigor: you can define a falsifiable question, build the right measurement, control confounders, interpret noisy results, and communicate uncertainty honestly.
The engineering ability to navigate an unfamiliar codebase, build reliable research infrastructure, and turn a promising idea into a working system.
Clear written and verbal communication and the ability to collaborate across research, systems, and product disciplines in a fast-moving environment.
Research experience in agentic reinforcement learning, post-training, preference learning, reward or verifier modeling, evaluation, or environment design.
Experience with large-model training, distributed inference, high-throughput rollout systems, or performance-sensitive ML infrastructure.
Notable publications, open-source contributions, datasets, benchmarks, or research artifacts that other people use.
Domain expertise in coding agents, mathematical reasoning, scientific discovery, tool use, long-horizon planning, or computer interaction.
Experience transferring a research result into a production model, product, or dependable shared system.
Mission…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).