×
Register Here to Apply for Jobs or Post Jobs. X

Model Research Engineer

Job in Zürich, 8058, Zurich, Kanton Zürich, Switzerland
Listing for: Rapidata AG
Full Time position
Listed on 2026-07-25
Job specializations:
  • IT/Tech
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Data Scientist, AI Evaluation
  • Research/Development
    Data Scientist, AI Evaluation
Salary/Wage Range or Industry Benchmark: 120000 - 180000 CHF Yearly CHF 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Reward Model Research Engineer
Location: Zürich

About Rapidata

Rapidata provides an API to humans that is revolutionizing the data generation and annotation industry. We deliver highly scalable, extremely fast human feedback that fuels the AI systems of the future, powering RLHF and DPO training data collection at internet speed for frontier AI labs. Our network reaches over 20 million active annotators across 192 countries, distributing micro-tasks ("Rapids") and returning verified labels in near real-time.

The Role

We re looking for a Reward Model Research Engineer to design, train, and product ionize the reward models work alongside Rapidata s real-time human feedback for high-quality training signal for RLHF and DPO. This is a research-to-production role: you ll work on the modeling problems that determine whether noisy, large-scale human preference data becomes a reliable reward signal, including process reward models (PRMs) that score multi-step reasoning.

You ll be closing the loop between our human-in-the-loop data collection infrastructure and the reward models that consume it and accompany it, tackling generalization across generator models, robustness to annotator noise, and inference-time guidance techniques, then validating your models in a live system operating at global scale.

What You ll Do
  • Design, train, and evaluate reward models, including process reward models (PRMs), from large-scale human preference data

  • Build and improve inference-time guidance methods to make reward-guided generation more robust

  • Develop training setups that improve generalization across diverse generator/policy models

  • Translate reward modeling research into production-ready pipelines that plug directly into Rapidata s RLHF/DPO data flywheel

  • Collaborate with the data platform team to design data collection and active learning strategies that reduce reward model training bottlenecks

  • Rigorously benchmark reward models for accuracy, robustness, and reliability before deployment

  • Communicate findings clearly to both technical and non-technical stakeholders, including partner AI labs

What We re Looking For
  • Hands-on research experience building reward models for LLMs or diffusion models, ideally through an MSc/PhD thesis or equivalent applied research

  • Solid understanding of reinforcement learning fundamentals and RLHF/DPO training pipelines

  • Practical experience with inference-time guidance techniques and evaluating multi-step reasoning

  • Strong Python and deep learning framework skills (PyTorch)

  • Experience turning research prototypes into validated, production-ready systems

  • Solid statistical and mathematical foundation

  • Excellent English communication skills, both oral and written

Nice-to-Have
  • Experience with agentic system safety, guardrails, or LLM-based tool-calling agents

  • Publications or open-source contributions in reward modeling, reasoning, or reinforcement learning

  • Experience with hierarchical or model-based RL

  • Based in or willing to relocate to Zürich

What We Offer
  • Competitive salary and equity in a startup with strong growth, IP, and backing from top-tier VCs

  • Opportunity to join a fast-growing startup early, with an outsized opportunity to shape where the company goes

  • Opportunities for personal and professional growth as our team expands

  • Fun and open (startup) culture

  • Spacious mountain-view office near Sihlcity, Zürich, with terrace, table tennis, pizza oven, hammock, and BBQ

  • Hardware budget tailored to your preferences

  • Unlimited snacks and drinks of your choice

#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary