×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer – GPU​/HPC Infrastructure; Remote – MENA)

Remote / Online - Candidates ideally in
UAE/Dubai
Listing for: Saturn Cloud
Remote/Work from Home position
Listed on 2026-10-02
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 360000 - 600000 AED Yearly AED 360000.00 600000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)

This is a fully remote role open to candidates across the MENA region

About Saturn Cloud

Saturn Cloud builds infrastructure for running AI, machine learning, and data workloads  platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.

We’re looking for a GPU/HPC-focused Site Reliability Engineer based in the MENA region to help operate and troubleshoot the large-scale GPU infrastructure supporting Saturn Cloud Token Factory.

This role is focused on the infrastructure below and around the Kubernetes layer. You’ll serve as a technical escalation point for GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU performance issues affecting production inference workloads.

The ideal candidate comes from GPU infrastructure, HPC, AI infrastructure, neocloud, hyperscaler, or large-scale ML platform operations rather than traditional application support.

We value engineers who actively use modern AI coding agents to improve the speed and quality of their engineering and operational work.

What You’ll Do
  • Diagnose and resolve production issues affecting large-scale GPU inference infrastructure
  • Troubleshoot NVIDIA datacenter GPUs, drivers, CUDA compatibility, and GPU container runtimes
  • Investigate GPU health issues, including Xid errors and hardware or driver failure modes
  • Diagnose PCIe, NUMA, GPU placement, NVLink, and NVSwitch issues
  • Troubleshoot multi-GPU and multi-node workloads
  • Investigate high-performance networking and distributed communication failures
  • Distinguish application and inference-runtime issues from GPU, fabric, topology, node, driver, or hardware failures
  • Work with Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins
  • Use production observability and GPU metrics to diagnose reliability and performance issues
  • Work directly with GPU-cloud and infrastructure providers when incidents require hardware or fabric investigation
  • Produce clear technical evidence showing where the failure occurs and what the appropriate infrastructure team needs to investigate
What We’re Looking For Core Skills
  • Hands-on experience using AI coding agents and agentic development tools to accelerate engineering, debugging, automation, and operational workflows
  • NVIDIA datacenter GPU administration
  • NVIDIA driver installation, upgrades, and troubleshooting
  • Strong understanding of CUDA and driver compatibility
  • Experience with NVML and nvidia-smi
  • GPU health diagnostics, including Xid errors and common hardware/driver failure modes
  • Understanding of PCIe topology, NUMA, and GPU placement
  • Understanding of NVLink and NVSwitch fundamentals
  • Familiarity with Kubernetes GPU Operator and device plugins
  • Experience running containerized GPU workloads
High-Performance Networking

You should have strong experience with several of the following:

  • Infini Band
  • RDMA
  • RoCE
  • NCCL
  • GPUDirect RDMA
  • NCCL testing and distributed workload diagnostics
  • Bandwidth and latency troubleshooting

You should be able to determine whether a production issue originates in the application, GPU, network fabric, topology, node, or another underlying infrastructure layer.

Inference Infrastructure

You don’t need to be an ML researcher, but you should understand how modern inference workloads exercise GPU infrastructure.

  • vLLM, NVIDIA Dynamo, Triton, or comparable inference runtimes
  • Tensor and pipeline parallelism
  • Model loading and GPU memory consumptionKV cache
  • Continuous batching
  • GPU and memory utilization
  • Out-of-memory diagnosis
  • Multi-GPU and multi-node inference
  • Basic inference performance analysis, including throughput, latency, and GPU saturation
Additional Skills
  • Kubernetes troubleshooting sufficient to independently investigate GPU workloads inside a cluster
  • Prometheus/Grafana and NVIDIA…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary