SWE-Bench Task Auditor
Northern, Floyd County, Kentucky, USA
Listed on 2026-09-03
-
Software Development
Software Testing, Python, AI QA / Validation Engineer
About Open Train
Open Train is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a profile that reflects their experience, and apply in minutes.
As an Open Train contractor, you can develop a durable portfolio of AI training work while finding opportunities that match your technical background. Creating an Open Train account is free.
About AI Training WorkAI training is the human side of building artificial intelligence. Technical contributors review code, assess model outputs, and evaluate benchmark tasks so AI systems can become more accurate, reliable, and useful.
This work puts experienced software professionals close to the development of cutting-edge AI systems. Remote projects can offer flexible schedules and the opportunity to apply practical engineering judgment to advanced model evaluation.
The RoleOpen Train is recruiting a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.
The role combines practical software-engineering judgment with careful evaluation of whether benchmark tasks are correct, reproducible, and resistant to answer leakage or reward hacking. You will provide concise feedback grounded in defined evaluation criteria.
- Remote contract role for candidates located in the United States
- Approximately 40 hours per week, with a stated commitment of 20 or more hours weekly
- Compensation of $70 to $90 per hour
- Part-time contractor engagement
You will review software-engineering tasks at the repository level and assess whether they accurately represent the intended work. Your findings will help identify benchmark weaknesses that could distort AI model evaluation.
- Review repository-level tasks for quality, correctness, and reproducibility
- Audit reference patches and determine whether they properly address the intended task
- Inspect test runners, grading behavior, and containerized execution for reliability
- Assess Docker isolation and overall grading integrity
- Identify answer leakage, reward hacking, and other evaluation weaknesses
- Write concise, rubric-based feedback describing findings and recommended improvements
This role requires at least three years of professional software-engineering experience, along with meaningful open-source contribution or maintainer experience. The listing is marked entry level, but candidates must meet the stated professional experience and technical requirements.
- At least three years of professional software-engineering experience
- Open-source contribution or maintainer experience, including merged pull requests or committer responsibilities
- Ability to audit reference patches, test runners, Docker isolation, and grading integrity
- Fluency in Python and at least one of Java, Go, Type Script, or C++
- Judgment in detecting answer leakage and reward hacking in software-engineering evaluations
- Familiarity with SWE-Bench Verified or similar repository-level benchmarks is helpful
- Prior code-review or task-grading experience is helpful
- Maintainer history on major Python open-source projects such as Django, Flask, scikit-learn, sympy, or pytest is valuable
This opportunity is suited to software engineers who can move comfortably between source code, test infrastructure, containerized execution, and evaluation criteria. It may be especially relevant to open-source contributors and maintainers who understand how repository-level changes should be tested and reviewed.
Strong candidates will be able to explain technical findings clearly, distinguish legitimate task difficulty from benchmark defects, and recognize when evaluation setups create opportunities for leakage or reward hacking.
- Professional software engineers with strong repository-level debugging judgment
- Open-source maintainers and contributors with merged pull requests or committer responsibilities
- Developers experienced with Python and another listed programming language
- Engineers familiar with benchmark evaluation, code review, or task grading
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).