×
Register Here to Apply for Jobs or Post Jobs. X

AI Software Engineer – LLM Evaluation & Automation; Remote

Remote / Online - Candidates ideally in
Washington, District of Columbia, 20022, USA
Listing for: Stage 4 Solutions Inc
Remote/Work from Home position
Listed on 2026-10-02
Job specializations:
  • Software Development
    AI QA / Validation Engineer, Software Testing, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 72 - 78.57 USD Hourly USD 72.00 78.57 HOUR
Job Description & How to Apply Below
Position: AI Software Engineer – LLM Evaluation & Automation (Remote)
AI Software Engineer LLM Evaluation & Automation (Remote)

We are looking for an AI Software Engineer for a B2B high-tech company. In this role, you will build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.

This is a 6-month (extensions likely), 40-hour/week, remote role in the US.

This is a W2 role as a Stage 4 Solutions employee. Health benefits and 401K are offered.

Must be able to work Pacific Time (PST) hours.

Responsibilities
  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees), so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and efficiency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.
Requirements:
  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar language.
  • Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
  • Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
Preferred:
  • Hands-on work with AI-powered coding tools and agentic applications, such as Claude Code, Devin, or Open Code.
  • Experience designing benchmarks or evaluations for software systems, especially execution based grading that verifies against tests.
  • Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them against human raters.
  • Experience with build-system-aware test selection, such as Bazel or mapping changed files to the tests that cover them.
  • Experience building reproducible test environments and managing versioned evaluation datasets. Comfortable writing up methodology and results for engineering leadership.

Stage 4 Solutions is an equal opportunity employer. We celebrate diversity and are committed to providing employees with an inclusive environment that is free of discrimination and harassment. All employment decisions are based on the job requirements and candidates' qualifications, without regard to race, color, religion/belief, national origin, gender identity, age, disability, marital status, genetic information or other applicable legally protected characteristics.

Compensation:

$72/hr.

- $78.57/hr. on W2

#LI-SW1

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary