More jobs:
Research Engineer, Research Scientist, AI Systems Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-08-16
Listing for:
Jobtailor
Full Time
position Listed on 2026-08-16
Job specializations:
-
Research/Development
AI Evaluation, Research Scientist
Job Description & How to Apply Below
- Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution
- Turn real research workflows and model failures into data and evaluation flywheels
- Improve model research capabilities through agent harnesses, synthetic data, RL environments, and model training
- Build and maintain safe, reliable integrations between models and OpenAI’s research infrastructure
- Develop research agents, experiment-orchestration systems, and sandboxed runtimes supporting real research workflows
- Create metrics and economic models to understand RSI’s effects on research productivity, model capabilities, and safety of internal deployments
- Work across research, engineering, product, and infrastructure to automate research workflows and improve research productivity
- Contribute across the full lifecycle of model training, evaluation, and deployment
- Research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems
- Strong generalist ability to move between open-ended research and practical implementation
- Ability to turn ambiguous problems into clear results
- Effective collaboration across systems, data, model training, evaluations, and research teams
- Experience building and maintaining data pipelines, tooling, and infrastructure for emerging AI capabilities
- Comfort working on problems without clear definitions or established playbooks
- Rigorous thinking about scientific quality, research taste, safety, privacy, reliability, performance, and scale
- Excitement about using increasingly capable AI systems to accelerate meaningful research
Demonstrates expertise in LLM Training, Model Evaluations, and Research Infrastructure, with a strong ability to automate research workflows and enhance research productivity through effective collaboration and rigorous scientific thinking.
Highest-signal resume keywords- LLM Training
- Model Evaluations
- Agent Systems
- Data Pipeline Development
- Research Infrastructure
- Model Training
- Experiment Design
- Synthetic Data Generation
- Research Workflow Automation
- Evaluation Metrics Development
- Effective Collaboration
- Rigorous Thinking
- Problem-Solving
- Large-Scale Distributed Systems
- AI Capabilities
- Research Productivity
- Model Capabilities
- Safety and Privacy
- Research Infrastructure
- Experiment-Orchestration Systems
- Sandboxed Runtimes
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×