×
Register Here to Apply for Jobs or Post Jobs. X

Cluster Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: STN Inc
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

Position Summary

We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.

Responsibilities
  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.

  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.

  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.

  • Build and support production AI infrastructure running hundreds to thousands of GPUs.

  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.

  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.

  • Design and optimize GPU networking using Infini Band or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.

  • Configure and tune distributed AI software stacks including:

    • Py Torch
    • NCCL
    • CUDA
    • UCX
    • MPI
    • Slurm
    • Pyxis/Enroot
  • Optimize GPU scheduling and resource allocation for both training and inference environments.

  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.

  • Identify performance regressions and troubleshoot distributed training issues at scale.

  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.

  • Work closely with ML engineers to improve training scalability and inference efficiency.

  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.

  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.

Required Qualifications
  • 7+ years designing or operating large-scale Linux infrastructure.
  • 5+ years supporting production GPU clusters for AI or HPC workloads.
  • Demonstrated experience building multi-node GPU training environments from the ground up.
  • Deep expertise with distributed PyTorch training.
  • Extensive experience troubleshooting and optimizing NCCL communications.
  • Strong understanding of distributed AI communication patterns, including:
    • All Reduce
    • Reduce Scatter
    • All Gather
    • Broadcast
    • Point-to-point communications
  • Experience benchmarking distributed training using tools such as:
    • nccl-tests
    • NVIDIA DCGM
    • Nsight Systems
    • MLPerf (preferred)
  • Strong understanding of GPU memory management, including:
    • KV Cache
    • Activation checkpointing
    • Tensor Parallelism
    • Pipeline Parallelism
    • Data Parallelism
  • Experience optimizing LLM inference throughput, including:
    • Tokens/sec optimization
    • Batch sizing
    • Continuous batching
    • KV cache tuning
    • Memory bandwidth optimization
  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
  • Expert-level Linux systems administration skills.
  • Experience with Slurm workload manager.
  • Experience using Pyxis and Enroot for containerized GPU workloads.
  • Strong scripting skills using Python and Bash.
Technical Expertise AI Frameworks
  • Py Torch
  • CUDA
  • NCCL
  • Triton (preferred)
  • TensorRT-LLM (preferred)
Cluster Scheduling
  • Slurm
  • Pyxis
  • Enroot
GPU Networking
  • Infini Band
  • RoCE v2
  • RDMA
  • GPUDirect RDMA
  • GPUDirect Storage
  • UCX
  • MPI
  • Network topology optimization
  • Congestion control
  • QoS
  • ECN/PFC
  • High-speed Ethernet (200/400/800 GbE)
Storage
  • Parallel file…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary