Researcher,Agent Post-Training,API & Power-Users Job San Francisco area,California USA,Software Development

Researcher, Agent Post-Training, API & Power-Users

Agents - San Francisco

About the Team

The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve.

We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste.

Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use.

About the Role

As a member of the API & power-users team, you will improve the capabilities, reliability, and product fit of OpenAI’s agentic models for power users and API developers. You might design evals from real developer workflows, build training environments around production-like tool use, turn qualitative model failures into training data, evals, or post-training interventions, or drive a behavior improvement from discovery through post-training, integration, and launch.

This role is intentionally broad. The strongest candidates are comfortable turning ambiguous model behavior problems into concrete progress, whether that means improving tool use, planning, instruction following, recovery from mistakes, or how models behave in API-based workflows.

You will work closely with researchers, engineers, API/product teams, Codex, infrastructure, and safety/alignment partners to decide which behaviors matter, how to measure them, how to train them, and when they are ready for major model runs. This is a high-agency role for people who want their work to show up directly in frontier models used by expert users and developers.

In this role, you might

Design and run experiments that improve model behavior in API and power-user workflows: function calling, tool use, coding, planning, long-horizon execution, factuality, instruction following, error recovery, and calibrated reasoning.
Build evals, graders, and environments from real developer and power-user workflows, then turn observed failures into training data, model-behavior hypotheses, and shipped improvements.
Partner with API and power-users to identify high-leverage behavior gaps and convert product signals into post-training interventions.
Improve how models behave when composed into systems: using tools reliably, respecting developer intent, handling partial failures, asking for clarification when appropriate, and maintaining coherence across multi-step tasks.
Own end-to-end model behavior projects, from qualitative failure analysis through data generation, training experiments, eval design, integration into major runs, and launch readiness.
Develop feedback loops that use power-user traces, API usage patterns, and production-like environments to discover the next frontier of agentic model failures and gaps.
Help decide which agentic capabilities, behavioral fixes, and partner-team integrations are ready for inclusion in major model runs.
Debug hard failures in shipped or near-shipped models by moving between traces, evals, training data, model outputs, and product context.
Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior.
Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments.

You might thrive in this…