×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer - AI Infrastructure

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Hamilton Barnes Associates Limited
Full Time position
Listed on 2026-07-13
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 350000 USD Yearly USD 350000.00 YEAR
Job Description & How to Apply Below
Position: Staff Site Reliability Engineer - AI Infrastructure

Looking for a role with plenty of growth opportunities?

Join a rapidly growing AI infrastructure provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The organization partners with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads.

This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering deployment, GPU health, distributed training, networking, and incident response. Working as a senior technical leader, the role focuses on improving infrastructure reliability, operational standards, and platform performance across complex AI environments.

Ready to make a move? Get in touch and apply today!

Responsibilities
  • Lead high-priority incident response across distributed GPU infrastructure environments
  • Diagnose and resolve issues across the stack, including PyTorch, NCCL, CUDA, drivers, networking fabrics, and hardware layers
  • Own day-to-day operational health of large-scale GPU fleets, including lifecycle management, validation, firmware rollouts, upgrades, and repair workflows
  • Build and maintain observability systems, GPU telemetry platforms, automated remediation tooling, and health-check frameworks
  • Define and scale operational practices, including on-call rotations, escalation processes, incident response, and postmortem standards
  • Partner closely with infrastructure, product, and platform engineering teams to improve reliability and scalability
  • Participate in customer-facing technical discussions, incident reviews, architecture workshops, and workload planning sessions
  • Influence physical infrastructure design, including rack layouts, power density, cooling strategies, burn-in procedures, and network topology decisions
  • Mentor engineers across reliability engineering, systems operations, and incident management practices
  • Contribute to a long-term reliability strategy for hyperscale AI infrastructure environments
Skills/Must Have
  • Multiple years of hands-on experience operating large-scale GPU infrastructure environments
  • Proven Staff-level SRE or infrastructure engineering experience supporting mission-critical production systems
  • Deep expertise with NVIDIA GPU platforms including H100, H200, B200, or GB200 systems
  • Strong understanding of GPU memory hierarchy, ECC behaviour, NVLink, NVSwitch, thermal management, and hardware failure analysis
  • Production experience with Infini Band, RoCE, and high-performance distributed training fabrics
  • Deep understanding of distributed AI training technologies, including NCCL, CUDA, PyTorch Distributed, FSDP, Deep Speed, and Megatron
  • Strong software engineering skills in Go, Python, or Rust
  • Experience building production-grade automation, tooling, fleet management systems, or reliability platforms
  • Hands-on experience with Kubernetes GPU environments, Slurm, or HPC schedulers
  • Strong Linux systems expertise, including kernel tuning, CUDA lifecycle management, cgroups/name spaces, BPF/performance analysis, and firmware operations
  • Calm, structured incident response capabilities within high-pressure production environments
  • Ability to communicate effectively with highly technical customers, providers, and executive stakeholders
Desirable Skills
  • Experience building custom GPU fleet health systems or fabric controllers
  • Expertise with distributed storage systems such as VAST, Weka, Lustre, or GPFS
  • Experience optimizing distributed training efficiency, checkpointing, and multi-thousand-GPU job performance
  • Background supporting enterprise AI infrastructure customers in customer-facing technical roles
  • Open-source contributions within the GPU, Kubernetes, or AI infrastructure ecosystem
  • Public speaking, technical writing, or community leadership within AI infrastructure or HPC domain
Benefits
  • Huge stock options
  • Company bonus
  • Unlimited PTO
  • 401K + 4% match
Salary
  • $350,000 base salary
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary