AI Benchmark Quality Engineer Task Evaluation
Listed on 2026-09-22
-
Quality Assurance - QA/QC
Quality Engineering, IT QA Tester / Automation
Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level tasks, reference patches, test harnesses, and grading integrity, and provide rubric-based written feedback.
Responsibilities include identifying gaps in tooling, ensuring task reproducibility across environments, and delivering concrete, rubric-based assessments to guide model training and
The following opening is for a AI Benchmark Quality Engineer for Task Evaluation with Obsidian.
Full responsibilities and requirements are described in the listing above.
Learn more about the AI Benchmark Quality Engineer for Task Evaluation role in the description above.
We appreciate your interest in this position.
Join Obsidian and contribute to our ongoing work.
Take a moment to read everything above and see whether this role is right for you.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).