×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer (SRE

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Recruiting From Scratch
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 170000 - 250000 USD Yearly USD 170000.00 250000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE)

Site Reliability Engineer (SRE)

Location:

San Francisco, CA / Palo Alto, CA Company Stage of Funding:
Growth-Stage AI Infrastructure Company ($80M Raised) Office Type:
Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity

We're representing a rapidly growing AI infrastructure company building a next-generation GPU cloud platform for enterprises, startups, and AI researchers. Their platform provides flexible access to GPU compute through intelligent reservation, marketplace, and consumption models that help customers optimize performance, availability, and cost.

Backed by Sequoia Capital and Lightspeed with more than $80 million in funding, the company has achieved 6x revenue growth over the past year. As demand for AI infrastructure accelerates, they're investing heavily in reliability engineering to build the automation, observability, and platform infrastructure that powers their multi-cloud GPU marketplace at scale.

What You Will Do
  • Design, build, and own the observability platform supporting a large-scale, multi-cloud GPU infrastructure.
  • Develop monitoring, distributed tracing, dashboards, and alerting systems using modern observability tooling.
  • Define and implement SLIs, SLOs, and operational metrics across customer-facing APIs and internal platform services.
  • Build automation that eliminates repetitive operational work and improves platform reliability.
  • Develop production tooling in Python or Go for infrastructure management, health checks, reconciliation, and capacity optimization.
  • Design and maintain Infrastructure-as-Code using Terraform, Pulumi, and Kubernetes.
  • Improve platform resiliency through incident response, root cause analysis, and long-term reliability improvements.
  • Partner closely with Platform, Product, and Engineering teams to ensure new services are designed for operational excellence.
  • Help establish infrastructure engineering standards, reliability practices, and operational processes as the company scales.
  • Participate in production on-call rotations while continuously reducing operational burden through automation.
Ideal Background
  • 3–10 years of experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or Platform Engineering.
  • Strong experience building production automation and operational tooling rather than solely responding to incidents.
  • Proven experience designing and operating large-scale Kubernetes environments.
  • Strong cloud infrastructure experience across AWS, GCP, Azure, or multi-cloud environments.
  • Experience designing distributed systems with a strong understanding of networking fundamentals.
  • Proficiency with Python and/or Go for building production-grade infrastructure tooling.
  • Experience implementing observability platforms using Prometheus, Grafana, Open Telemetry, or similar technologies.
  • Strong understanding of Linux systems, containers, Docker, and production operations.
  • Excellent communication skills with the ability to collaborate across engineering teams.
Compensation and Benefits
  • Base salary: $170,000–$250,000.
  • Competitive equity package.
  • Visa transfer sponsorship available.
  • Four-day onsite schedule across San Francisco and Palo Alto offices (all engineers collaborate in Palo Alto on Mondays).
  • Opportunity to help define the reliability and operational foundation of one of the fastest-growing AI infrastructure platforms.
  • Significant ownership over observability, automation, and production infrastructure.
  • Work alongside experienced engineers solving large-scale distributed systems and cloud infrastructure challenges.
  • Join a high-growth, venture-backed company building the infrastructure powering the next generation of AI applications.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary