×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Member of Technical Staff (Language Model Evaluations

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Artificial Analysis, Inc.
Full Time position
Listed on 2026-08-01
Job specializations:
  • Research/Development
    AI Evaluation
Salary/Wage Range or Industry Benchmark: 170000 - 230000 USD Yearly USD 170000.00 230000.00 YEAR
Job Description & How to Apply Below
Position: Member of Technical Staff (Language Model Evaluations)

Job Description – Member of Technical Staff (Language Model Evaluations)

Location:

San Francisco (preferred), Sydney, Melbourne, Brisbane

About Artificial Analysis

Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don’t just measure the cutting edge of AI, they are actively shaping the frontier.

Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.

We are a team of 40+, on track to double by end of year, backed by Nat Friedman (Git Hub, Meta), Daniel Gross (SSI, Meta), Andrew Ng (Google Brain, Deep Learning.ai, Amazon), Adam D’Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders.

The Opportunity

Language model evaluation is the sharpest question in AI: what can these systems actually do? Our answers, from the Artificial Analysis Intelligence Index to AA-Omniscience, AA-Briefcase and our coding agent evaluations, are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them.

This is a role for people who want to build frontier benchmarks: designing evaluations that stay ahead of frontier capabilities, constructing datasets that resist contamination, and measuring what everyone else has not yet worked out how to measure. You will run your work across every major model as it releases and publish results the whole industry reads.

The center of the role is building. Analysis and lab collaboration wrap around the evaluation work, with our commercial team owning client relationships day to day.

What You’ll Do
  • Design Next-Generation Frontier Evals: Conceive and ship the next generation of frontier evaluations, like AA-Briefcase and AA-Omniscience, across reasoning, knowledge, coding, agentic capability and beyond
  • Build Evaluation Datasets and Infrastructure: Construct the datasets, harnesses and scoring systems behind our benchmarks, engineered for contamination resistance and repeatability at frontier scale
  • Shape the Future Intelligence Index: The evaluations you build will contribute to future versions of the Artificial Analysis Intelligence Index and other areas of our platform, defining how the industry measures frontier capability
  • Publish Influential Analysis: Produce the reports, indexes and data visualizations that shape how the industry understands language model progress
  • Work with Frontier Labs on Pre-Release Models: Benchmark the leading labs’ systems, including pre-release and newly launched models, working directly with their research teams; our commercial team owns client relationships day to day, so your time stays on the science
  • Evaluate Every Major Model: Run our evaluation suite across frontier releases as they land, and own the integrity of the results the industry quotes
  • Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking
What We’re Looking For

You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail.

Backgrounds include: evaluation and benchmarking teams at AI labs; research or engineering roles at evaluation-focused organizations; ML engineers who have built evaluation harnesses and datasets in production; or academic researchers in NLP and ML evaluation with a strong record of published work.

Required:
  • 3+ years of relevant professional experience, across industry or research
  • Strong analytical and critical thinking skills
  • Strong Python, with hands‑on experience running evaluation harnesses and building datasets
  • Deep familiarity with the LLM evaluation landscape: the major benchmarks and their failure modes, contamination,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary