×
Register Here to Apply for Jobs or Post Jobs. X

AI Research Scientist, Learning & Evaluation

Job in Beverly Hills, Los Angeles County, California, 90211, USA
Listing for: Studyfetch
Full Time position
Listed on 2026-07-31
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

AI Research Scientist, Learning & Evaluation

Studyfetch Beverly Hills, California, United States

About this position

About Studyfetch

Study Fetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We’re growing fast with backing from top-tier investors and a mission that’s redefining the future of education and ethical learning.

Why this role exists

We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.

Nobody has settled how to measure whether an AI tutor teaches. Public benchmarks tell you a model can answer a question. They don't tell you whether a fourteen-year-old understood the explanation, whether the model handed over the answer when it should have asked a follow-up, or whether the student could still do the problem a week later. We train and tune our own models for learning outcomes rather than leader board scores, and that only works if someone can define what better means and prove when we've hit it.

That's this role. You'll build the evaluation and measurement layer that sits under model training, product decisions, and learning science across both Study Fetch and Honen. This is a founding-team role. You'll work directly with the people making the decisions, and the standard you set is the one every model and feature gets held to.

What we believe

Every learner deserves the chance to succeed. Study Fetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.

Accessible to everyone.

Meet people where they are.

Learning never stops.

We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.

What you'll own

Evaluation for the models we train.

We fine-tune our own model family for tutoring, and you decide how we know whether a new checkpoint is better than the last one. That covers accuracy and reasoning, and it also covers child safety, resistance to sycophancy, and whether the model teaches Socratically instead of answering outright. You'll design the evals, run them against every candidate model, and hold the release bar.

An internal benchmark for multi-turn tutoring.

Single-turn Q&A benchmarks miss almost everything we care about. You'll build and maintain our benchmark for real tutoring conversations, decide what it measures, and defend those choices to researchers outside the company. Expect to publish parts of it.

The link between product data and model training.

The Learn Engine records what worked for past students at the personal, course, topic, and global level. You'll turn that into training signal and into evidence: which interventions moved mastery, which ones only moved engagement, and which model behaviors correlate with a student actually learning.

Data quality for expert-verified content.

We build assessment questions with subject-matter experts, starting in nursing licensure and expanding into medical and legal. You'll measure agreement between experts, catch where the official answer key is out of date, and design how those verifications feed back into training.

Analytics across both products.

Retention, activation, feature adoption, conversion, and how each of those moves when a model changes. You'll build the dashboards and reporting that product and leadership actually use, and you'll say plainly when the data can't answer the question yet.

Instrumentation we don't have.

You'll find the missing events, telemetry, and logging, then work with engineering to add them. Most of the interesting questions here are currently unanswerable because nobody logged the right thing.

The measurement bar for the team.

How we run…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary