Distributed Computing DevOps Engineer; IT-CE-LCG--LD
Listed on 2026-07-15
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Support
Location: Genf
Company Description
At CERN, the European Organization for Nuclear Research, physicists and engineers are probing the fundamental structure of the universe. Using the world's largest and most complex scientific instruments, they study the basic constituents of matter—fundamental particles that collide together at close to the speed of light. This process gives physicists clues about how particles interact and provides insights into the fundamental laws of nature.
Job DescriptionAre you a motivated Computing Engineer looking to contribute to a global scientific infrastructure? Join CERN’s IT Computing for Experiments group (IT-CE), responsible for activities supporting the CERN physics community, from coordination of core functions and services of the Worldwide LHC Computing Grid (WLCG) to development of advanced computing solutions and co‑design of future computing models with the experiments. As part of the CE‑LCG section, you will help shape the future of distributed computing within WLCG, a global collaboration of more than 170 computing centres in over 40 countries supporting cutting‑edge particle physics research.
As part of the High‑Luminosity LHC (HL‑LHC) programme, you will contribute to the operation and evolution of a large‑scale distributed computing infrastructure, ensuring reliable data processing and transfer services for global scientific collaborations. You will work in a stimulating international and multidisciplinary environment, contributing directly to the computing projects and operations of one or more scientific collaborations, as part of a team of computing engineers and physicists driving continuous improvement of services and operational procedures.
Responsibilities- Play a key role in the development, integration, and automation of operational tools and workflows in preparation for the HL‑LHC.
- As part of your experiment liaison role, contribute directly to selected computing projects and operations of one or more scientific collaborations.
- Monitor and support the operation of a high‑scale global distributed computing infrastructure across multiple international sites.
- Evolve and operate monitoring, accounting, and topology systems used to track service availability, data transfers, and resource utilisation.
- Support daily operational activities, including incident tracking, debugging of service issues, and coordination with external sites and users.
- Contribute to the integration and adoption of AI‑driven tools and automation technologies within operational processes to improve operational efficiency, automate routine tasks, and enhance service performance and reliability.
- Master’s Degree or PhD or equivalent relevant experience in Physics, Computing, Engineering or a related field.
- Experience in Linux‑based distributed computing environments.
- Proficiency in Python programming and scripting languages.
- Experience with web frameworks, relational and non‑relational databases.
- Experience with Dev Ops workflows including Code Management (Jira, Git Lab), Continuous Integration, Delivery and Deployment, Containerisation & Orchestration (Kubernetes), and tools widely used for monitoring (Prometheus, Grafana).
- Capacity to critically interpret and analyse data.
- Experience working with large‑scale systems such as cloud or grid infrastructures.
- Experience with the computing of large scientific projects or working with scientific communities would be considered a strong asset.
- Ability to work in an international and collaborative environment.
- Knowledge of programming techniques and languages.
- Knowledge of system configuration tools.
- Installation, operation and maintenance (preventive and corrective) of computing systems.
- Development of application software.
- Testing, diagnosing and optimisation of software.
- Development of data visualisation applications.
- Knowledge and application of software life‑cycle tools and procedures.
- Knowledge of operating systems.
- Working in Teams: gaining trust and collaboration from others.
- Solving Problems: identifying, defining and assessing problems, taking action to address them.
- Communicating Effectively:…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: