Remote STEM Researcher - AI Evaluation & Benchmark Design
Washington, District of Columbia, 20022, USA
Listed on 2026-10-03
-
Research/Development
AI Evaluation, Data Annotation/ AI Labeling, Research Scientist
Weekday 1 is seeking researchers to develop rigorous evaluation benchmarks for frontier AI models. The role focuses on transforming experimental design and data-driven analysis into sophisticated, multi-step benchmarks that challenge current AI systems.
You will work with AI researchers to uncover subtle reasoning errors and methodological flaws while contributing to the continuous refinement of evaluation methodologies. This is a fully remote, full-time engagement with about 35 hours per week.
The following role is for a Remote STEM Researcher - AI Evaluation & Benchmark Design with Weekday 1.
This Full Time position is for the Remote STEM Researcher - AI Evaluation & Benchmark Design role at Weekday 1.
We are seeking a motivated Remote STEM Researcher - AI Evaluation & Benchmark Design to join Weekday 1 in United States.
Consider building your career as a Remote STEM Researcher - AI Evaluation & Benchmark Design at Weekday 1.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).