Sr. Site Reliability Engineer
Listed on 2026-07-27
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Overview
One of Insight Global’s customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern Dev Ops practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support deployments, build reliable systems, and strengthen platform capabilities across Kubernetes, AWS, and CI/CD pipelines. They will own the observability strategy across a growing application ecosystem.
This person will partner closely with product and engineering teams, helping drive reliability, monitoring, logging, and platform scalability initiatives. This is a rolling contract and will be 5 days a week onsite in Irvine, CA.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances.
If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
- Provide reliability engineering and platform engineering support for deployments and system design across Kubernetes, AWS, and CI/CD pipelines.
- Own the observability strategy and drive monitoring, logging, and metrics across a growing application ecosystem.
- Collaborate with product and engineering teams to improve platform reliability, performance, and scalability.
- Partner with cross-functional teams to implement best practices for security, resilience, and incident response.
- 3–6+ years of Site Reliability Engineering, Dev Ops, or Platform Engineering experience
- Strong AWS experience (GCP and Azure are a plus)
- Hands-on Kubernetes platform engineering (multi cluster, deployment, troubleshooting); experience building and maintaining CI/CD pipelines
- Observability and monitoring experience within modern application environments
- Experience with Prometheus and Grafana for metrics, monitoring, and logging
- Excellent communication and stakeholder management skills
- Python and/or React development experience
- Experience with Kafka or event streaming platforms
- Neo4j experience
- AI Infrastructure experience
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).