Site Reliability Engineer @Phoenix, AZ
Listed on 2026-08-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Site Reliability Engineer
Role is onsite based out of Scottsdale, AZ. Years of experience required: 7+. Required skills include service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud). Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys. Experience working with programming languages such as Go, Python, Java, Rust etc. Working knowledge on with one or more databases
- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases. Experience in transitioning platforms to the cloud and Containerization – GCP and Rancher. Experience maintaining containerized app in GKE/RKE/AKE environments. Experience implementing cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution. Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc.). Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.
Preferred skills include proven experience managing application availability, building creative solutions to manage repetitive activities, improving gating and detect for applications at every touchpoint for a 24 x 7 high availability platform exposed to critical clients and customers. Working knowledge of monitoring tools - Splunk, App-dynamics, grafana/Prometheus and Dynatrace. Experience with tools like Rally, Confluence and other CI/CD extenders. Hands-on experience with implementing in-memory caching solutions.
Experience on Redis DB is a plus. Excellent debugging skills across variety of integrated technical platforms on API gateway. Hands-on with GCS, Cloud SQL, Spanner and Firestore. Extensive experience in enterprise level infrastructure and operations. Experience in high availability and distributed systems, Linux and Windows administration, troubleshooting and support. Monitor and troubleshoot Hashi Corp Vault environments, ensuring minimal downtime and rapid recovery from incidents.
Working knowledge on Vertex AI, Gen AI and Bigquery. Google Cloud Platform (GCP) Containerization, Kubernetes. Infrastructure as Code (Terraform), CI/CD (Git Hub Actions), and Helm. Automation and scripting using Python, Ansible, and Node.js. Monitoring and observability with
Years of
Experience:
14.00 Years of Experience
Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).