Site Reliability Engineer II
Listed on 2026-08-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, Cybersecurity
Site Reliability Engineer II
Location:
Chandler, Arizona (Hybrid)
Duration: 12 months
Role Overview
This position is for a Site Reliability Engineer responsible for the reliability and support of on-premise and external cloud Container Platforms, including Azure, AWS, and Google. The role involves monitoring, troubleshooting, and enhancing the performance and security of container environments such as Open Shift, Rancher (RKE), and Azure (AKS). The ideal candidate will be a key stakeholder in the design of cloud services and will work to improve automation and operational excellence.
Key Responsibilities
- Provide reliability and support for Container Platforms on-premise and in external clouds (Azure, AWS, Google).
- Monitor and troubleshoot performance, connectivity, and security issues for Open Shift, Rancher (RKE), and Azure (AKS) environments.
- Conduct deep dives into systemic and latent reliability issues, managing incidents and problems.
- Identify, analyze, and resolve infrastructure vulnerabilities and application deployment issues.
- Perform blameless Root Cause Analysis (RCA) and partner with engineering and operation teams to implement fixes.
- Manage application onboarding and provide troubleshooting support throughout the application lifecycle.
- Identify and implement automation opportunities to reduce operational tasks and improve efficiency.
- Collaborate with risk and compliance teams to implement controls and remediate vulnerabilities.
- Ensure resiliency during implementation and work with engineering teams to resolve resiliency problems.
- Participate in a 24x7 on-call rotation following a follow-the-sun model.
Required Qualifications
Education:
B.S./M.S. degree in Computer Science or a related technical field, or equivalent practical experience.
Experience:
Minimum of 5+ years of hands-on experience supporting Kubernetes, Open Shift, RKE, or EKS Container platforms.
Technical
Skills:
- Experience with Python, Ansible, Golang, and shell scripting.
- Experience with major services related to Compute, Storage, Network, and Security.
- Experience with monitoring tools like Prometheus and Dynatrace, and cloud-native tools like Azure Monitor and Log Analytics.
- Strong understanding of complex IAM infrastructure, including Active Directory, Azure AD, and SSO solutions like Ping Identity.
- Advanced knowledge of Linux OS, DNS, DHCP, Kerberos, and Windows Authentication.
- Experience with CI/CD tools such as Git and Jenkins, and Git Ops models.
- Excellent understanding of Linux/Windows operating systems administration.
- Experience in container security and vulnerability remediation.
Preferred Qualifications
- Experience in Open Shift, RKE, and CSP Kubernetes services such as AKS and EKS.
- Experience in Terraform, ArgoCD, Tekton, and K-native technologies.
- Experience in agile deployment methodologies (Git Ops).
- Knowledge of various container runtimes.
- Familiarity with the operator deployment pattern.
- Experience working in a highly available multi-datacenter environment.
- Experience with monitoring tools such as Prometheus, Splunk, Dynatrace, or Sysdig.
- Understanding of cost management, inventory management, and the Fin Ops model.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).