AI Trainer – CHEMISTRY
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-10-08
-
Science
AI Evaluation, Research Scientist
is looking for experienced scientists and engineers to evaluate how frontier AI models handle real technical work: analyzing test, measurement or process data, sizing and verifying a design, designing an experiment and reading it out, interpreting a simulation, writing the technical report. You bring the judgment you have built catching the unit error in a test file, the multiple comparisons trap in a process improvement study, and the assumption that does not hold at the boundary.
We bring the model output that judgment is needed to grade.
In this role, you will design challenging, realistic tasks drawn from your own practice, such as a calculation package with stated assumptions and checks, a test plan and data analysis, a failure or deviation investigation, a design trade study, an experimental protocol with acceptance criteria, a circuit or system design review, or a technical report, run them through frontier AI agents, and evaluate what comes back against a professional standard.
Across our STEM and Data Science programs, tasks are grounded in real day-to-day workflows and checked against frontier models so only genuinely hard tasks make it through. Some projects are authored and verified in code, so coding / scientific computing is a strong plus and is required on those projects.
You will work with realistic professional files, the kind a practitioner in your field actually handles, which you assemble yourself. Some tasks are compact, built around a handful of files; others are larger scenarios that take several days to build. In every case the goal is the same: a task a competent professional in your field would complete correctly and a frontier model currently gets wrong.
This is not a traditional science or engineering role. You will be helping build better AI by putting your knowledge to work in a structured, flexible, fully remote environment. The work is long-form and self-directed, and clear written reasoning matters as much as technical depth.
Responsibilities- ? Design challenging, realistic technical tasks drawn from your own day-to-day work: the scenario, a prompt phrased the way you would brief a trusted colleague, and the supporting files an engineer or scientist would need (test data, drawings or schematics, specifications, simulation outputs, lab records, reports), which you author yourself.
- ? Run those tasks through frontier AI models and evaluate the deliverable they produce (the calculation, analysis, design review or report) against the standard you would hold a colleague to.
- ? Compare two model outputs on identical prompts and files, decide which performed better, and document where each fell short.
- ? Write detailed grading rubrics that specify what a correct deliverable must contain (the right assumptions stated, the right method, the right magnitudes and units, the right failure modes considered), and explain in writing why a response passes or fails each one.
- ? Flag concrete failures with evidence: unit and scaling errors, misread data, unsupported conclusions, fabricated or ignored source files, missed safety or boundary conditions, and off-brief interpretation of the ask.
- ? Contribute across your discipline and adjacent ones, and review and refine tasks built by other experts.
- ? 2+ years of applied experience preferred in one of the core sciences (mathematics, physics, chemistry, or biology) or in an engineering discipline (electrical and electronics, civil and structural, materials, chemical and process, environmental and earth sciences, or similar) or in data science.
- ? In progress Bachelor’s degree or higher. We prefer 2+ years of experience outside undergraduate study.
- ? Coding / scientific computing is a strong plus and is required on some projects: comfortable writing and verifying work in code (Python or similar) and working at the command line.
- ? Hands-on with real data and the tools of your field: instrument and test data, measurement files, simulation outputs, schematics or drawings, and the analysis tools that go with them (MATLAB, Python or R; SPICE and PCB tools; FEA or CFD; GIS; chromatography or spectroscopy software).
- ? Comfort with statistical and experimental reasoning: experiment design, measurement error and uncertainty, and the common statistical traps. The assessment leans on this.
- ? Working understanding of adjacent sub-disciplines, enough to assess work outside your own specialty and point out what was done correctly or incorrectly.
- ? Hands-on…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).