×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Evlo AI
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 200000 USD Yearly USD 140000.00 200000.00 YEAR
Job Description & How to Apply Below

About The Role

The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.

About

The Role

The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.

The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs.

Key Responsibilities
  • Design, build, and maintain production infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and Pulumi
  • Implement comprehensive observability stacks using Prometheus, Grafana, Open Telemetry, and Datadog for real-time monitoring and alerting
  • Drive incident response and conduct thorough blameless post-mortems to continuously improve system resilience and prevent recurrence
  • Automate deployment pipelines and release engineering processes using Git Hub Actions, ArgoCD, and Kubernetes
  • Optimize cloud infrastructure costs, resource utilization, and database performance without compromising system reliability
  • Establish and enforce security compliance, IAM policies, and disaster recovery runbooks across all production environments
What We Are Looking For
  • 3–7 years of experience in Site Reliability Engineering, Dev Ops, or systems engineering roles within high-growth tech environments
  • Deep expertise in Kubernetes administration, containerization (Docker), and service mesh architectures
  • Strong proficiency in scripting languages such as Python or Go for automation and tool development
  • Hands-on experience with modern CI/CD pipelines, Git Ops workflows, and infrastructure as code frameworks
  • Solid understanding of networking fundamentals (TCP/IP, DNS, TLS, load balancing) and distributed systems failure modes
  • Bonus:
    Experience with Chaos Engineering practices, eBPF for observability, or holding professional cloud certifications (AWS/GCP)
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary