Site Reliability Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-09-04
Listing for:
Evlo AI
Full Time
position Listed on 2026-09-04
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
About The Role
The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.
AboutThe Role
The role owns the reliability, scalability, and performance of production distributed systems handling massive traffic scale across multi-region cloud environments.
The team works closely with software engineering squads to embed resilience into architecture, automate operational toil, and maintain strict SLAs.
Key Responsibilities- Design, build, and maintain production infrastructure on AWS or GCP using Infrastructure as Code tools such as Terraform and Pulumi
- Implement comprehensive observability stacks using Prometheus, Grafana, Open Telemetry, and Datadog for real-time monitoring and alerting
- Drive incident response and conduct thorough blameless post-mortems to continuously improve system resilience and prevent recurrence
- Automate deployment pipelines and release engineering processes using Git Hub Actions, ArgoCD, and Kubernetes
- Optimize cloud infrastructure costs, resource utilization, and database performance without compromising system reliability
- Establish and enforce security compliance, IAM policies, and disaster recovery runbooks across all production environments
- 3–7 years of experience in Site Reliability Engineering, Dev Ops, or systems engineering roles within high-growth tech environments
- Deep expertise in Kubernetes administration, containerization (Docker), and service mesh architectures
- Strong proficiency in scripting languages such as Python or Go for automation and tool development
- Hands-on experience with modern CI/CD pipelines, Git Ops workflows, and infrastructure as code frameworks
- Solid understanding of networking fundamentals (TCP/IP, DNS, TLS, load balancing) and distributed systems failure modes
- Bonus:
Experience with Chaos Engineering practices, eBPF for observability, or holding professional cloud certifications (AWS/GCP)
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×