Site Reliability Engineer – GPU/HPC Infrastructure; Remote – MENA)
UAE/Dubai
Listed on 2026-10-02
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
This is a fully remote role open to candidates across the MENA region
About Saturn CloudSaturn Cloud builds infrastructure for running AI, machine learning, and data workloads platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.
We’re looking for a GPU/HPC-focused Site Reliability Engineer based in the MENA region to help operate and troubleshoot the large-scale GPU infrastructure supporting Saturn Cloud Token Factory.
This role is focused on the infrastructure below and around the Kubernetes layer. You’ll serve as a technical escalation point for GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU performance issues affecting production inference workloads.
The ideal candidate comes from GPU infrastructure, HPC, AI infrastructure, neocloud, hyperscaler, or large-scale ML platform operations rather than traditional application support.
We value engineers who actively use modern AI coding agents to improve the speed and quality of their engineering and operational work.
What You’ll Do- Diagnose and resolve production issues affecting large-scale GPU inference infrastructure
- Troubleshoot NVIDIA datacenter GPUs, drivers, CUDA compatibility, and GPU container runtimes
- Investigate GPU health issues, including Xid errors and hardware or driver failure modes
- Diagnose PCIe, NUMA, GPU placement, NVLink, and NVSwitch issues
- Troubleshoot multi-GPU and multi-node workloads
- Investigate high-performance networking and distributed communication failures
- Distinguish application and inference-runtime issues from GPU, fabric, topology, node, driver, or hardware failures
- Work with Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins
- Use production observability and GPU metrics to diagnose reliability and performance issues
- Work directly with GPU-cloud and infrastructure providers when incidents require hardware or fabric investigation
- Produce clear technical evidence showing where the failure occurs and what the appropriate infrastructure team needs to investigate
- Hands-on experience using AI coding agents and agentic development tools to accelerate engineering, debugging, automation, and operational workflows
- NVIDIA datacenter GPU administration
- NVIDIA driver installation, upgrades, and troubleshooting
- Strong understanding of CUDA and driver compatibility
- Experience with NVML and nvidia-smi
- GPU health diagnostics, including Xid errors and common hardware/driver failure modes
- Understanding of PCIe topology, NUMA, and GPU placement
- Understanding of NVLink and NVSwitch fundamentals
- Familiarity with Kubernetes GPU Operator and device plugins
- Experience running containerized GPU workloads
You should have strong experience with several of the following:
- Infini Band
- RDMA
- RoCE
- NCCL
- GPUDirect RDMA
- NCCL testing and distributed workload diagnostics
- Bandwidth and latency troubleshooting
You should be able to determine whether a production issue originates in the application, GPU, network fabric, topology, node, or another underlying infrastructure layer.
Inference InfrastructureYou don’t need to be an ML researcher, but you should understand how modern inference workloads exercise GPU infrastructure.
- vLLM, NVIDIA Dynamo, Triton, or comparable inference runtimes
- Tensor and pipeline parallelism
- Model loading and GPU memory consumptionKV cache
- Continuous batching
- GPU and memory utilization
- Out-of-memory diagnosis
- Multi-GPU and multi-node inference
- Basic inference performance analysis, including throughput, latency, and GPU saturation
- Kubernetes troubleshooting sufficient to independently investigate GPU workloads inside a cluster
- Prometheus/Grafana and NVIDIA…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).