×
Register Here to Apply for Jobs or Post Jobs. X

Apertus Engineer: Evaluations

Job in Zürich, 8058, Zurich, Kanton Zürich, Switzerland
Listing for: ETH Zürich
Full Time position
Listed on 2026-07-19
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 120000 - 180000 CHF Yearly CHF 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Apertus Engineer: Evaluations 100%
Location: Zürich

We are seeking a skilled engineer to join the Apertus evaluation effort. The ideal candidate will build and operate the evaluation codebase and pipelines that inform our training and release decisions, keeping results consistent between training and serving. This role requires strong Python engineering, hands‑on LLM evaluation experience, and the ability to work collaboratively in a research-focused environment.

Project background

We train open foundation models with hundreds of billions of parameters on thousands of GPUs on one of the largest AI‑ready supercomputers in Europe. The team counts more than a dozen full‑time engineers working alongside leading researchers from EPFL and ETH Zürich, has released the Apertus 1 and Apertus 1.5 models, and works with over thirty academic collaborators to deliver fully open (open source), responsibly trained, multilingual, multimodal AI models for research and industry.

Apertus is trained and developed on Alps, the Swiss National Supercomputing Centre’s (CSCS) supercomputing infrastructure. The role requires someone who is comfortable working in an HPC environment and collaborating with researchers and infrastructure engineers.

Job description

The engineer will own the evaluation codebase and pipelines that inform training decisions and releases.

Evaluation infrastructure
  • Build and maintain the evaluation codebase and pipelines for Apertus models, from checkpoints during training to released models
  • Make evaluations run quickly and at scale: parallel execution on Alps, efficient use of inference backends, caching, and result tracking
  • Reduce mismatch between evaluations during training and during serving: consistent tokenisation, chat templates, prompting, and sampling across evaluation harnesses and inference engines
  • Debug evaluation failures, regressions, and inconsistencies across backends
Benchmark coverage
  • Integrate and run the evaluations the project cares about. The design of new evaluations is owned by collaborating researchers and engineers; this role makes them run reliably and at scale
  • Cover image and audio evaluations alongside text within the same pipeline
  • Integrate new benchmarks as the field evolves, working with our academic collaborators to onboard the benchmarks they create, and validate that metrics and harness implementations are trustworthy
Comparative and third‑party evaluation
  • Evaluate third‑party services and other open and closed models against the same benchmark suite, producing directly comparable and reproducible results
  • Provide evaluation results, reports, and dashboards that support training decisions (data mixtures, ablations) and release decisions
  • Work closely with the engineers focused on safety, deployment, and community needs, and integrate the evaluations they create into the shared pipeline
Profile

Essential

  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field. Exceptional
  • BSc candidates with strong engineering experience will also be considered
  • Strong Python and software engineering skills, including experience building robust data or evaluation pipelines
  • Experience with LLM evaluation: established harnesses (e.g. lm-evaluation-harness) or custom benchmark tooling
  • Strong collaboration and communication skills and ability to work across research and engineering teams
  • Prior hands‑on experience in the core domains of this role is required. This can be project or study based experience; formal work experience is preferred
  • A high degree of flexibility: priorities, tools, and day‑to‑day tasks shift with training schedules, releases, and a fast‑moving field
  • Experience running evaluations at scale on GPU clusters (Slurm or similar) and with inference engines such as vLLM or SGLang
  • Familiarity with agentic evaluation and agentic harnesses: tool use, sandboxed execution environments, benchmarks such as SWE‑bench or similar
  • Experience with multimodal (image or audio) model evaluation

Strongly preferred

  • An eye for statistical rigor: variance across runs, prompt sensitivity, significance of differences between models

Nice to have

  • Published research in the domains relevant to this role, or familiarity with…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary