×
Register Here to Apply for Jobs or Post Jobs. X

Principal Research Scientist - Evaluations

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Canva
Full Time position
Listed on 2026-09-01
Job specializations:
  • Research/Development
    AI Evaluation, Research Scientist
Job Description & How to Apply Below

Job Description

Join the team redefining how the world experiences design.

Hey, g'day, mabuhay, kia ora,, hallo, vítejte!

Thanks for stopping by. We know job hunting can be a little time consuming and you're probably keen to find out what's on offer, so we'll get straight to the point.

Where and how you can work

Our head office is in Sydney, Australia, but San Francisco is now home to our US operations. The role is listed as hybrid, meaning we are incredibly flexible and empower you to work where you prefer - whether that's at home or at the office.

About the role

Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships.

We are looking for a Principal Research Scientist who defines what evaluation needs to become as the space gets harder, rather than running the playbook we already have. You will own how Canva evaluates generative quality across the whole of Canva Research, including problems we have not framed yet: new modalities, evaluation that reflects real differences between content types, user segments and markets, and a much tighter link between what our metrics say and what users and the business actually experience.

This is a Canva-wide craft leadership role, setting direction across our research groups in Australia, Europe, the US and China. You will be the person others come to when the numbers and the eyes disagree.

What you'll own

The evaluation strategy for Canva Research. Define what great evaluation looks like across design, image, video, audio and agentic workflows, and set the long-term direction for how Canva measures generative quality as the space evolves. You will shape the principles teams use to trade off evaluation compute, human data spend and signal fidelity, and you will defend them.

The science of measurement itself. Auto-raters and MLLM judges are only as good as their correlation with the thing they proxy for. You will treat that correlation as a research problem: validating metrics against human preference and downstream product outcomes, quantifying judge bias, and catching benchmark saturation and contamination before they quietly stop telling us anything.

The link between evaluation and what actually matters. Evaluation should reflect user experience and product impact, not just perform well in isolation. You own closing that gap, including the fact that a good evaluator is not automatically a good reward model for RL. That extends to the full experience rather than the artefact alone, editability included.

One standard across every region. Our teams in London, Vienna, San Francisco, Sydney and China all need to know they are measuring the same thing. You will build the shared evaluation layer that makes results comparable, and you will spot the collaboration opportunities nobody has been chartered to own yet. Taking that initiative is a core expectation of this role, not a bonus.

Focus areas

Human preference at scale. Rubric design, rater guidelines, inter-rater reliability, and calibration across markets and cultural contexts. Aesthetic judgement is not universal, and our evaluation systems need to hold that honestly rather than average it away.

Learned quality models and reward signals. Reward modelling and preference learning that feed post-training (RLHF, RLAIF) and inference-time selection, and being explicit about where a good judge does not translate into a good reward model.

Vision-Language Models for quality understanding. Novel architectures and training approaches for models that understand what makes a design effective, not just well-formed. These become the reward signal and feedback loop for our design generation models, so their failure modes are our failure modes.

Agentic and automated evaluation. MLLM-as-a-judge systems, model-based grading, and the infrastructure to run hundreds of evaluations against…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary