Evaluation Engineer: AI Coding Benchmarks & Tests
Listed on 2026-10-02
-
Software Development
AI QA / Validation Engineer, Software Testing, AI Reliability/ Performance Engineer
Mercor is seeking software engineers to build and refine evaluations for frontier models. You’ll transform completed PRs into engineering tasks and use our in‑house framework to benchmark across models such as Claude Code and Codex.
You’ll collaborate with an in‑house research team, write prompts and tests, and help refine grading criteria. Strong programming skills, experience with large codebases, and clear communication are essential to own task development and iteration.
The Evaluation Engineer: AI Coding Benchmarks & Tests role at Mercor is now open for applications in United States.
Take a moment to read everything above and see whether this role is right for you.
This posting is for the Evaluation Engineer: AI Coding Benchmarks & Tests role at Mercor, based in United States.
We are looking to fill the Evaluation Engineer: AI Coding Benchmarks & Tests position at Mercor in United States.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).