GenAI Evaluation Scientist - LLM Benchmarks & Failures
Listed on 2026-10-07
-
Research/Development
AI Evaluation, Data Scientist
Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities.
Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence
Are you ready to take on the GenAI Evaluation Scientist - LLM Benchmarks & Failures role at Scale?
This opportunity is part of our work in Bio & Pharmacology & Health, Other.
The advertised compensation is 166..
We aim to respond to suitable candidates as soon as possible.
Full responsibilities and requirements are described in the listing above.
Learn more about the GenAI Evaluation Scientist - LLM Benchmarks & Failures role in the description above.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).