×
Register Here to Apply for Jobs or Post Jobs. X

Apertus Engineer: Infrastructure

Job in 1001, Lausanne, Canton de Vaud, Switzerland
Listing for: Eidgenössische Technische Hochschule Zürich
Full Time position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    Systems Engineer, Unix/Linux
Salary/Wage Range or Industry Benchmark: 120000 - 160000 CHF Yearly CHF 120000.00 160000.00 YEAR
Job Description & How to Apply Below

We are seeking a skilled infrastructure engineer to join the Apertus team. The ideal candidate will own the container image stack behind our pre‑training, post‑training, and serving workloads, and collaborate with CSCS engineers to keep large‑scale training on Alps stable and fast. This role requires strong Linux and container skills, experience with HPC environments, and the ability to work collaboratively across research, engineering, and operations teams.

Project

background

The Apertus project, a joint effort between EPFL, ETH Zürich, and CSCS, is seeking a practical and motivated infrastructure engineer to help build the next version of Apertus. The successful candidate will own the container image stack and work closely with CSCS to keep large‑scale training stable and fast.

We train open foundation models with hundreds of billions of parameters on thousands of GPUs on one of the largest AI‑ready supercomputers in Europe. The team counts more than a dozen full‑time engineers working alongside leading researchers from EPFL and ETH Zürich, has released the Apertus 1 and Apertus 1.5 models, and works with over thirty academic collaborators to deliver fully open, responsibly trained, multilingual, multimodal AI models for research and industry.

Apertus is trained and developed on Alps, the Swiss National Supercomputing Centre's (CSCS) supercomputing infrastructure. The role requires someone who is comfortable working in an HPC environment and collaborating with researchers and infrastructure engineers.

The engineer will enable stability and throughput for the Apertus pre‑training and post‑training pipelines by maintaining the ML system images and partnering with CSCS on the underlying infrastructure.

ML system image maintenance
  • Build, maintain, and upgrade container images for all core ML development phases: pre‑training, post‑training/alignment, and model serving/deployment
  • Target the ARM‑based (aarch
    64, Grace‑Hopper) node architecture of Alps, managing the full dependency stack (CUDA, NCCL, PyTorch, training and serving frameworks)
  • Keep image builds reproducible, versioned, and documented, including CI for builds and upgrades
  • Validate images against reference pre‑training and post‑training workloads together with Apertus engineers, and maintain working launch examples
Compute partnership and efficiency
  • Serve as the primary technical point of contact with CSCS engineers and researchers regarding compute, reliability, and efficiency
  • Work collaboratively with CSCS staff to identify and implement improvements in the efficiency and performance of the underlying compute infrastructure, overlapping with ML systems performance engineering
  • Contribute to systemic improvements in CSCS‑based resources (network, storage, scheduling) relevant to large‑scale LLM training
  • Document and disseminate institutional knowledge about the CSCS infrastructure and best practices for leveraging these high‑performance systems
Infrastructure stress testing
  • Stress test the infrastructure using representative pre‑training and post‑training workloads, building on the project's existing recipes and examples, to validate stability and throughput after image upgrades, system maintenance, and configuration changes
  • Work closely with Apertus pre‑training and post‑training engineers to debug cluster‑level problems affecting stability and throughput: node failures, networking, storage performance, checkpointing, and scheduling
  • Support the Apertus serving stack, which builds on the same images (operation of the serving stack is owned by a separate engineer)
Profile
  • MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field
  • Exceptional BSc candidates with strong engineering experience will also be considered
  • Hands‑on experience with HPC environments: job schedulers such as Slurm, shared file systems, and multi‑node GPU systems
  • Strong Linux systems and container skills (Docker/Podman and HPC runtimes such as enroot or Apptainer)
  • Strong collaboration and communication skills and ability to work across research, engineering, and operations teams
  • Prior hands‑on experience in the core domains of this role is…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary