×
Register Here to Apply for Jobs or Post Jobs. X

SWE-Bench Task Auditor

Remote / Online - Candidates ideally in
Northern, Floyd County, Kentucky, USA
Listing for: OpenTrain AI
Full Time, Remote/Work from Home position
Listed on 2026-09-03
Job specializations:
  • Software Development
    Software Testing, Python, AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 70 - 90 USD Hourly USD 70.00 90.00 HOUR
Job Description & How to Apply Below
Location: Northern

About Open Train

Open Train is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a profile that reflects their experience, and apply in minutes.

As an Open Train contractor, you can develop a durable portfolio of AI training work while finding opportunities that match your technical background. Creating an Open Train account is free.

About AI Training Work

AI training is the human side of building artificial intelligence. Technical contributors review code, assess model outputs, and evaluate benchmark tasks so AI systems can become more accurate, reliable, and useful.

This work puts experienced software professionals close to the development of cutting-edge AI systems. Remote projects can offer flexible schedules and the opportunity to apply practical engineering judgment to advanced model evaluation.

The Role

Open Train is recruiting a SWE-Bench Task Auditor to evaluate repository-level software-engineering benchmark tasks used to train and assess AI models. You will examine task specifications, reference patches, test harnesses, Docker isolation, and grading integrity.

The role combines practical software-engineering judgment with careful evaluation of whether benchmark tasks are correct, reproducible, and resistant to answer leakage or reward hacking. You will provide concise feedback grounded in defined evaluation criteria.

  • Remote contract role for candidates located in the United States
  • Approximately 40 hours per week, with a stated commitment of 20 or more hours weekly
  • Compensation of $70 to $90 per hour
  • Part-time contractor engagement
What You'll Do

You will review software-engineering tasks at the repository level and assess whether they accurately represent the intended work. Your findings will help identify benchmark weaknesses that could distort AI model evaluation.

  • Review repository-level tasks for quality, correctness, and reproducibility
  • Audit reference patches and determine whether they properly address the intended task
  • Inspect test runners, grading behavior, and containerized execution for reliability
  • Assess Docker isolation and overall grading integrity
  • Identify answer leakage, reward hacking, and other evaluation weaknesses
  • Write concise, rubric-based feedback describing findings and recommended improvements
Required Qualifications

This role requires at least three years of professional software-engineering experience, along with meaningful open-source contribution or maintainer experience. The listing is marked entry level, but candidates must meet the stated professional experience and technical requirements.

  • At least three years of professional software-engineering experience
  • Open-source contribution or maintainer experience, including merged pull requests or committer responsibilities
  • Ability to audit reference patches, test runners, Docker isolation, and grading integrity
  • Fluency in Python and at least one of Java, Go, Type Script, or C++
  • Judgment in detecting answer leakage and reward hacking in software-engineering evaluations
  • Familiarity with SWE-Bench Verified or similar repository-level benchmarks is helpful
  • Prior code-review or task-grading experience is helpful
  • Maintainer history on major Python open-source projects such as Django, Flask, scikit-learn, sympy, or pytest is valuable
Who Should Apply

This opportunity is suited to software engineers who can move comfortably between source code, test infrastructure, containerized execution, and evaluation criteria. It may be especially relevant to open-source contributors and maintainers who understand how repository-level changes should be tested and reviewed.

Strong candidates will be able to explain technical findings clearly, distinguish legitimate task difficulty from benchmark defects, and recognize when evaluation setups create opportunities for leakage or reward hacking.

  • Professional software engineers with strong repository-level debugging judgment
  • Open-source maintainers and contributors with merged pull requests or committer responsibilities
  • Developers experienced with Python and another listed programming language
  • Engineers familiar with benchmark evaluation, code review, or task grading
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary