Apertus Engineer: Infrastructure
Listed on 2026-07-19
-
IT/Tech
Systems Engineer, Unix/Linux
Location: Zürich
We are seeking a skilled infrastructure engineer to join the Apertus team. The ideal candidate will own the container image stack behind our pre‑training, post‑training, and serving workloads, and collaborate with CSCS engineers to keep large‑scale training on Alps stable and fast. This role requires strong Linux and container skills, experience with HPC environments, and the ability to work collaboratively across research, engineering, and operations teams.
Projectbackground
The Apertus project, a joint effort between EPFL, ETH Zürich, and CSCS, is seeking a practical and motivated infrastructure engineer to help build the next version of Apertus. The successful candidate will own the container image stack and work closely with CSCS to keep large‑scale training stable and fast.
We train open foundation models with hundreds of billions of parameters on thousands of GPUs on one of the largest AI‑ready supercomputers in Europe. The team counts more than a dozen full‑time engineers working alongside leading researchers from EPFL and ETH Zürich, has released the Apertus 1 and Apertus 1.5 models, and works with over thirty academic collaborators to deliver fully open (open source), responsibly trained, multilingual, multimodal AI models for research and industry.
Apertus is trained and developed on Alps, the Swiss National Supercomputing Centre's (CSCS) supercomputing infrastructure. This role requires someone who is comfortable working in an HPC environment and collaborating with researchers and infrastructure engineers.
Job descriptionThe engineer will enable stability and throughput for the Apertus pre‑training and post‑training pipelines by maintaining the ML system images and partnering with CSCS on the underlying infrastructure.
ML system image maintenance- Build, maintain, and upgrade container images for all core ML development phases: pre‑training, post‑training/alignment, and model serving/deployment
- Target the ARM‑based (aarch
64, Grace‑Hopper) node architecture of Alps, managing the full dependency stack (CUDA, NCCL, PyTorch, training and serving frameworks) - Keep image builds reproducible, versioned, and documented, including CI for builds and upgrades
- Validate images against reference pre‑training and post‑training workloads together with Apertus engineers, and maintain working launch examples
- Serve as the primary technical point of contact with CSCS engineers and researchers regarding compute, reliability, and efficiency
- Work collaboratively with CSCS staff to identify and implement improvements in the efficiency and performance of the underlying compute infrastructure, overlapping with ML systems performance engineering
- Contribute to systemic improvements in CSCS‑based resources (network, storage, scheduling) relevant to large‑scale LLM training
- Document and disseminate institutional knowledge about the CSCS infrastructure and best practices for leveraging these high‑performance systems
- Stress test the infrastructure using representative pre‑training and post‑training workloads, building on the project's existing recipes and examples, to validate stability and throughput after image upgrades, system maintenance, and configuration changes
- Work closely with Apertus pre‑training and post‑training engineers to debug cluster‑level problems affecting stability and throughput: node failures, networking, storage performance, checkpointing, and scheduling
- Support the Apertus serving stack, which builds on the same images (operation of the serving stack is owned by a separate engineer)
- MSc or PhD in Computer Science, Data Science, Artificial Intelligence, Machine Learning, or a related field
- Exceptional BSc candidates with strong engineering experience will also be considered
- Hands‑on experience with HPC environments: job schedulers such as Slurm, shared file systems, and multi‑node GPU systems
- Strong Linux systems and container skills (Docker/Podman and HPC runtimes such as enroot or Apptainer)
- Strong collaboration and communication skills and ability to work across research, engineering, and operations teams
- Prior hands‑on…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: