Mathematics PhD - AI Evaluation Expert
Listed on 2026-08-28
-
Research/Development
AI Business & Operations, Mathematics, Research Scientist, Data Scientist
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code)
Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve.
Domains - depth required in at least two subdomains (with a coding focus)- Mathematics — numerical linear algebra, computational mechanics, computational finance
Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design
Write scientific prompts based on the input
Build the grading criteria that define a correct answer
Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed
PhD in mathematics, applied mathematics, computational mathematics, or a closely related field
Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance
Working proficiency in Python or R for scientific computing
Comfortable with Git/Git Hub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks
Publications in peer-reviewed journals
Prior scientific software or research engineering experience
Duration: 6 weeks
Commitment: part-time, 20+ hours per week
Start date:
immediate
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).