×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Research Engineer

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: SuperAnnotate AI
Full Time position
Listed on 2026-08-05
Job specializations:
  • Research/Development
    AI Evaluation
Job Description & How to Apply Below

Research Engineer

Super Annotate helps the world's leading AI teams build responsible, next-generation models powered by high-quality human data. We're a fast-growing Series B startup bridging the gap between advanced AI innovation and the data that drives it. Our global network of expert specialists, scalable managed operations, precise talent matching, and full project transparency ensure unmatched data quality sted by innovators like Databricks and Service Now - and backed by NVIDIA, Dell Technologies Capital, Databricks Ventures, Cox Enterprises, and Lionel Messi's Play Time VC - Super Annotate is proud to be the top-ranked AI data company on G2 for multiple consecutive years, including 2025.

The Impact You'll Make

Our research team is expanding to keep pace with a wave of frontier-facing work: internal research streams, client engagements that require real ML depth, and emerging opportunities at the cutting edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's relevant, and building out the process to reproduce and improve on it internally.

You'll own initiatives end to end: partnering with strategic project and technical leads to scope the work, building MVPs to validate ideas (including through human annotation and agents), and turning that work into something concrete – a customer dataset, a pilot, an internal dataset that becomes a paper or blog post, or a joint publication with a partner. You won't be handed a fully specified task list;

you'll be given a direction and the autonomy to turn it into a research plan.

This is a full-time, hybrid position based in San Francisco.

What You'll Do

  • Take a research direction and independently identify supporting resources – papers, benchmarks, blog posts – then implement or reimplement the relevant methods.
  • Build and own the process to reproduce prior work internally and identify ways to improve on it.
  • Own projects (for example, an RL/agentic environment build for a partner or a novel multimodal benchmark) end to end, including scoping, MVP implementation, and validation.
  • Partner with strategic project leads and technical leads to translate ambiguous requirements into a concrete, testable research plan.
  • Validate ideas through hands-on implementation, including annotating, evaluating, or sourcing data.
  • Turn research directions into tangible outputs – a paid customer dataset, a customer pilot, an internal dataset, or a paper/blog post for publication or conference presentation.
  • Bring an ML perspective to new opportunities — assessing technical feasibility of incoming requests and helping shape proposals where research depth is needed.

What You'll Bring

  • MS or PhD in ML, CS, or a related quantitative field – or equivalent demonstrated research experience (publications, significant open-source research work, industry research).
  • Real ML depth: you understand how models are trained and evaluated, not just how to call an API. You can read a paper, judge whether its claims hold, and reimplement the method.
  • Hands-on experience with at least one of: RL/agentic systems, AI/ML evaluation and benchmarking, or multimodal ML.
  • Strong Python and the engineering ability to build and ship your own experiments – eval harnesses, environments, infrastructure – without relying on a platform team.
  • High autonomy: you can turn an ambiguous direction into a concrete research plan and notice when something's off before being told.
  • Clear technical writing

Nice To Have

  • Publication track record (first-author preferred).
  • Experience with agent or multimodal benchmarks (OSWorld, MMMU, Web Arena, SWE-bench, or similar) or building RL environments/gyms.
  • Familiarity with reward modeling, reward hacking, or verifier/judge reliability.
  • Familiarity with synthetic data generation or human-in-the-loop (HITL) workflows.
  • Experience with cloud infrastructure and containerized environments.
  • A deep RL background specifically.

$180,000 - $280,000 a year In addition to the annual base salary, employees are eligible for an annual bonus paid out quarterly.

Why Super Annotate

This is a rare opportunity to…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary