×
Register Here to Apply for Jobs or Post Jobs. X

HPC Engineer - AI Infrastructure

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Hamilton Barnes Associates Limited
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 235000 - 315000 USD Yearly USD 235000.00 315000.00 YEAR
Job Description & How to Apply Below

Ready to take the next step in your career?

Join a VC-backed GPU cloud company at the founding stage, building the next generation of managed GPU compute and inference infrastructure. Backed by leading venture investors, the organization delivers fully managed AI infrastructure and offers the opportunity to help shape the technical foundation of a rapidly growing GPU cloud platform.

This company is seeking a Founding Engineer to take ownership of core GPU cloud infrastructure across Kubernetes, Slurm, GPU orchestration, distributed storage, high-bandwidth networking, and inference platforms. This hands-on role offers founding-level ownership, significant influence over platform architecture, and the opportunity to work closely with the founders to build and scale production AI infrastructure from the ground up.

Responsibilities:
  • Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
  • Build and manage GPU orchestration layers on top of core scheduling infrastructure
  • Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
  • Design and manage high-bandwidth networking supporting distributed training and inference at scale
  • Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
  • Build telemetry, observability, and automated remediation across the GPU fleet
  • Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
  • Own operational health, reliability, and performance of the platform end to end
  • Work directly with founders on architecture, roadmap, and technical strategy
  • Help define engineering culture, standards, and hiring as one of the first technical team members
Skills/Must Have:
  • 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production
  • Proven experience operating large-scale Kubernetes and Slurm clusters
  • Experience building and managing GPU orchestration layers on top of core schedulers
  • Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking
  • Deep knowledge of storage architectures for large-scale AI infrastructure
  • Comfort operating with founding-level ownership across the full infrastructure stack
  • Based in or willing to relocate to San Francisco
Benefits:
  • Founding engineer equity
  • Full benefits package
Salary:
  • $275,000 Base
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary