Site Reliability Engineering AI Evaluator
Listed on 2026-08-03
-
Software Development
Software Testing, AI QA / Validation Engineer, Software Engineer
Site Reliability Engineering AI Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.
Why this role mattersEngineering model quality lives or dies on whether the generated code actually compiles, passes tests, and handles edge cases. Aura One pairs experienced engineers with the modeling team to grade outputs the way a code reviewer would.
Responsibilities- Run and reproduce candidate code outputs in a sandboxed environment for Site Reliability Engineering AI Evaluator assignments.
- Grade cloud & platform engineering solutions for correctness, complexity, idiomatic style, and edge-case handling.
- Write minimal failing tests that demonstrate the bug a model output missed.
- Compare paired solutions and rank them with a written rationale tied to the rubric.
- Tag failure modes (compile error, runtime crash, off-by-one, security issue) with severity scores.
- Document recurring code-generation failures so the modeling team can target them.
- Calibrate against gold-standard reviews to keep inter-rater agreement above target.
- Strong day-job engineering experience — you can read, run, and debug unfamiliar code for Site Reliability Engineering AI Evaluator work.
- Comfort writing concise unit tests that capture a single failure mode.
- Familiarity with at least one of: pytest, Jest, JUnit, Go test, RSpec, or equivalent.
- Clear written reasoning — your review note has to convince another senior engineer.
- Reliable async availability for at least 10 hours per week.
- Prior code-review or technical-interview-grading experience is a plus.
- Reproduce a generated engineering solution to a coding task, run the test suite, and grade it.
- Write the smallest failing test that demonstrates a model's edge-case bug.
- Compare two paired solutions and rank them with a written rationale tied to the rubric.
- Triage a security issue surfaced by a model and document the patch the model should have produced.
- Open-source contributions or a public portfolio that demonstrates production-quality code.
- Experience with the target language's standard tooling, linters, and idiomatic style guides.
- Familiarity with security-review checklists (OWASP, CWE) and App Sec patterns.
- Code review
- Debugging
- Unit testing
- Software engineering judgment
- Cloud & platform engineering
- Software engineering and computer use
- AI evaluation
- Rubric writing
- Expert review
Remote — US-eligible. Remote
· Independent specialist contractor.
Employment type:
CONTRACTOR. Applicants must be authorized to work from US.
Hourly rate confirmed after the interview process.
Strong day-job engineering experience — you can read, run, and debug unfamiliar code for Site Reliability Engineering AI Evaluator work. Comfort writing concise unit tests that capture a single failure mode. Familiarity with at least one of: pytest, Jest, JUnit, Go test, RSpec, or equivalent. Clear written reasoning — your review note has to convince another senior engineer. Reliable async availability for at least 10 hours per week.
Prior code-review or technical-interview-grading experience is a plus.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).