Senior HPC Operations Engineer AI/ML Scale
Job Description & How to Apply Below
Core
42 is seeking a Senior HPC Operations Engineer to oversee daily HPC workloads powering AI/ML tasks. You will manage compute, storage, networking, and schedulers (Slurm/Kubernetes) across distributed environments, driving automation and reliability.
The role requires 7+ years in HPC/Dev Ops, strong GPU management, and expertise with Prometheus, Grafana, and DCGM. You will mentor engineers and contribute to on-call coverage while upholding security policies.
#J-18808-LjbffrPosition Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×