Site Reliability Engineer
Job in
Denver, Denver County, Colorado, 80285, USA
Listed on 2026-07-28
Listing for:
Evlo AI
Full Time
position Listed on 2026-07-28
Job specializations:
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
About The Role
The role focuses on scaling and maintaining the core cloud infrastructure that powers global production services. This position sits at the intersection of software engineering and systems engineering, ensuring that distributed systems are resilient, performant, and highly observable.
The team is responsible for architecting multi-region Kubernetes clusters, defining Infrastructure as Code standards, and building automated CI/CD deployment pipelines. The mission is to eliminate manual intervention through automation, mitigate production incidents, and maintain strict SLAs for millions of concurrent users.
Key Responsibilities- Design, provision, and manage multi-region cloud infrastructure using Terraform, Helm, and AWS/GCP cloud services
- Own the availability, latency, performance, efficiency, and capacity management of Kubernetes cluster environments
- Implement comprehensive observability stacks using Prometheus, Grafana, Open Telemetry, and Datadog to proactively detect and diagnose system degradation
- Develop internal tooling and automation in Go or Python to simplify deployment workflows and reduce operational toil
- Participate in a blameless post-mortem culture and share in an on-call rotation to quickly resolve production incidents
- Collaborate with product engineering teams to optimize application performance, containerize services, and architect fault-tolerant distributed systems
- 3–7 years of experience in SRE, Dev Ops, or systems engineering roles managing high-traffic production environments
- Strong hands‑on experience with container orchestration using Kubernetes (EKS, GKE, or self-managed)
- Deep proficiency in writing Infrastructure as Code using Terraform or Pulumi
- Solid software engineering foundation with strong coding skills in Python, Go, or Bash for automation
- Strong understanding of networking concepts (DNS, TCP/IP, VPC peering, load balancing) and Linux systems administration
- Bonus:
Experience with service meshes (Istio, Linkerd), Git Ops workflows (ArgoCD, Flux), or managing relational and No
SQL databases at scale
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×