More jobs:
Site Reliability Engineer
Job in
Scottsdale, Maricopa County, Arizona, 85251, USA
Listed on 2026-08-05
Listing for:
Apolis
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Network Engineer
Job Description & How to Apply Below
Site Reliability Engineer (SRE)
We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-availability applications in hybrid cloud environments. The ideal candidate will have strong expertise in cloud infrastructure, Kubernetes, automation, observability, container platforms, and application performance management. This role requires experience maintaining mission-critical enterprise systems, implementing reliability best practices, and driving automation across cloud-native platforms.
Role and Responsibilities:- Ensure the reliability, availability, and performance of large-scale enterprise applications.
- Develop automation scripts and dashboards for application performance monitoring and transaction tracking.
- Support hybrid infrastructure across on-premises and cloud environments.
- Manage containerized applications using Kubernetes (GKE/RKE/AKE) and Rancher.
- Implement cloud observability using Open Telemetry (OTEL), distributed tracing, and monitoring solutions.
- Automate infrastructure provisioning using Terraform, Helm, and CI/CD pipelines.
- Troubleshoot networking, application, infrastructure, and cloud platform issues.
- Administer cloud databases and caching solutions.
- Monitor and maintain Hashi Corp Vault environments.
- Collaborate with development, infrastructure, and operations teams to improve platform reliability and scalability.
- Support enterprise Infrastructure & Operations and 24x7 production environments.
- Site Reliability Engineering (SRE)
- Google Cloud Platform (GCP)
- Kubernetes (GKE/RKE/AKE)
- Rancher
- Docker
- Python
- Go
- Java
- Terraform
- Helm
- CI/CD
- Git Hub Actions
- Open Telemetry (OTEL)
- Splunk
- Grafana
- Prometheus
- Dynatrace
- App Dynamics
- Hashi Corp Vault
- GraphQL (Apollo/Prisma/Hasura)
- Redis
- Oracle
- PostgreSQL
- MongoDB
- SQL Server
- Cloud SQL
- Spanner
- Firestore
- Big Query
- Vertex AI
- GenAI
- Ansible
- Node.js
- TCP/IP
- HTTP
- DNS
- Load Balancing
- Service Mesh
- Linux
- Windows
- 7+ years of experience as a Site Reliability Engineer, Dev Ops Engineer, or Platform Engineer.
- Strong experience managing high-availability, cloud-native applications.
- Hands-on expertise with GCP, Kubernetes, Rancher, and container orchestration.
- Experience implementing Infrastructure as Code using Terraform and Helm.
- Strong scripting skills in Python, Go, Java, or similar languages.
- Experience with observability platforms including Open Telemetry, Splunk, Grafana, Prometheus, Dynatrace, or App Dynamics.
- Strong knowledge of networking, distributed systems, and cloud infrastructure.
- Excellent troubleshooting, automation, and communication skills.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×