Research Intern
Listed on 2026-10-09
-
Research/Development
AI Evaluation, Research Analyst, AI Business & Operations, Research Scientist
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You'll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
Aboutthe Role
As a Research Scientist Intern at Mercor, you'll work on research at the frontier of post-training, reinforcement learning with verifiable rewards (RLVR), data generation, and model evaluation.
You’ll investigate how datasets, rewards, and training methods affect the capabilities and behavior of large language models. This may include designing controlled experiments, developing new evaluation methodologies, conducting systematic failure analysis, and testing approaches to improve tool use, agentic behavior, and real-world reasoning.
You’ll work closely with research scientists, research engineers, and domain experts to turn open-ended questions into rigorous experiments. Your work will contribute to Mercor's research agenda and may support external publications, benchmark releases, and the development of frontier AI systems.
What You'll Do- Develop and investigate research questions related to post‑training, RLVR, data quality, and model evaluation.
- Design and run controlled experiments to understand how datasets, rewards, and training strategies affect model performance.
- Study reward‑shaping and post‑training methods, including approaches such as GRPO and DAPO.
- Develop methods for measuring data quality, usability, and performance uplift on key benchmarks.
- Design and evaluate datasets, rubrics, evaluators, and scoring frameworks for complex model capabilities.
- Conduct systematic error analysis to identify model failure modes and opportunities for improvement.
- Analyze experimental results and communicate findings through clear reports, research artifacts, and presentations.
- Build the research tooling and data pipelines needed to conduct experiments at scale.
- Collaborate with research scientists, research engineers, applied AI teams, and domain experts producing training and evaluation data.
- Contribute to research publications, benchmark releases, and other public research outputs where appropriate.
Currently pursuing a master's or PhD in computer science, machine learning, statistics, mathematics, or another relevant field.
- Demonstrated ability to formulate research questions, design experiments, and draw sound conclusions from empirical results.
- Demonstrated experience in at least one of the following:
- Training, fine‑tuning, or evaluating language models.
- Agentic AI system, RL environments.
- Developing benchmarks, evaluation methodologies, or data‑quality measures.
- At least one publication or open source project.
- Strong programming skills, particularly in Python, and the ability to write reliable research code.
- Familiarity with machine learning fundamentals, experimental design, and statistical analysis.
- Intellectual curiosity.
- Comfort operating in a fast‑paced research environment with…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).