LLM Evaluation Engineer - Benchmarks & Failure Analysis
Listed on 2026-09-30
-
Engineering
AI Evaluation -
Research/Development
AI Evaluation
Nous Research is seeking an engineer to strengthen our evaluation systems across lab work, benchmark design, and judge calibration. You will ship evaluation infrastructure that researchers rely on from day one, while owning end-to-end eval pipelines and failure analysis.
The role emphasizes high ownership on a small team, with opportunities to shape prompts, environments, and automated graders across GAIA-like benchmarks and new tasks.
This Full Time role is for the LLM Evaluation Engineer - Benchmarks & Failure Analysis role at Nous Research.
The position is based in United States.
This opportunity is part of our work in IT & Technology, Engineering.
The advertised compensation is 130..
We aim to respond to suitable candidates as soon as possible.
Full responsibilities and requirements are described in the listing above.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).