×
Register Here to Apply for Jobs or Post Jobs. X

GenAI Evaluation Scientist - LLM Benchmarks & Failures

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Scale
Full Time position
Listed on 2026-10-07
Job specializations:
  • Research/Development
    AI Evaluation, Data Scientist
Salary/Wage Range or Industry Benchmark: 166000 - 207000 USD Yearly USD 166000.00 207000.00 YEAR
Job Description & How to Apply Below

Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities.

Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence

Are you ready to take on the GenAI Evaluation Scientist - LLM Benchmarks & Failures role at Scale?

This opportunity is part of our work in Bio & Pharmacology & Health, Other.

The advertised compensation is 166..

We aim to respond to suitable candidates as soon as possible.

Full responsibilities and requirements are described in the listing above.

Learn more about the GenAI Evaluation Scientist - LLM Benchmarks & Failures role in the description above.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary