ML Engineer
Listed on 2026-10-04
-
Software Development
AI QA / Validation Engineer, Software Testing, AI Reliability/ Performance Engineer
ML Engineer is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real‑world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.
Category:
Coding, SWE & Agent Evaluation
· Pay: $130 / hr
·
Location:
Remote — US-eligible
· Contractor
ML Engineer is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real‑world correctness standards.
Role detailsAbout the role
ML Engineer is a remote engineering review track for evaluating production code, debugging traces, and developer‑facing AI outputs against real‑world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.
Engineering model quality lives or dies on whether the generated code actually compiles, passes tests, and handles edge cases. Aura One pairs experienced engineers with the modeling team to grade outputs the way a code reviewer would.
Judge generated code and software engineering agents. Read their debugging traces.
- Run and reproduce candidate code outputs in a sandboxed environment for ML Engineer assignments.
- Grade ml engineering solutions for correctness, style, and edge‑case handling.
- Write minimal failing tests that demonstrate the bug a model output missed.
- Compare paired solutions and rank them with a written rationale tied to the rubric.
- Tag failure modes (compile error, runtime crash, off‑by‑one, security issue) with severity scores.
- Strong day‑job engineering experience — you can read, run, and debug unfamiliar code for ML Engineer work.
- Comfort writing concise unit tests that capture a single failure mode.
- Familiarity with a testing framework. pytest, Jest, or JUnit. Go test, RSpec, or whatever you use.
- Clear written reasoning — your review note has to convince another senior engineer.
- Reliable async availability for at least 10 hours per week.
- Prior code‑review or technical‑interview‑grading experience is a plus.
Example tasks
- Reproduce a generated engineering solution to a coding task, run the test suite, and grade it.
- Write the smallest failing test that demonstrates a model's edge‑case bug.
- Compare two paired solutions and rank them with a written rationale tied to the rubric.
- Triage a security issue surfaced by a model and document the patch the model should have produced.
- Open-source contributions or a public portfolio that demonstrates production‑quality code.
- Experience with the target language's standard tooling, linters, and idiomatic style guides.
- Familiarity with security‑review checklists (OWASP, CWE) and App Sec patterns.
Hourly rate confirmed after the interview process.
Expected arrangement: contractor , with program‑defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation.
Skills used in matching- Code review
- Debugging
- Unit testing
- Software engineering judgment
- ML engineering
- AI model evaluation
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).