Senior Site Reliability Engineer
Job in
Chandler, Maricopa County, Arizona, 85249, USA
Listed on 2026-07-21
Listing for:
Koitecc Solutions
Full Time
position Listed on 2026-07-21
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Position Summary
The Senior Azure Site Reliability Engineer acts as an advanced senior individual contributor responsible for designing, implementing, and maturing reliability engineering capabilities across the enterprise Azure platform. The role focuses on complex technical problem solving, reliability architecture, automation strategy, observability maturity, platform resiliency, and operational excellence.
- Environment and Platform Scope
- Enterprise‑scale, governed Azure platform supporting infrastructure, platform services, and application workload enablement
- Primary tools and practices:
Microsoft Azure, Terraform / Terraform Enterprise, Azure Monitor, Azure Log Analytics, Dynatrace, CI/CD pipelines, Git, Python, Power Shell, Bash, Service Now/Jira or equivalent workflow tools - Core domain areas:
Azure landing zones, private networking, DNS, firewalls, private endpoints, observability, incident response, problem management, canary health checks, production readiness, and cloud governance - Role supports Azure platform reliability, operational readiness, observability, automation, and enterprise controls.
- Design solutions to visualize key production support metrics enabling Operational Readiness and Site Reliability Engineer teams to identify scenarios requiring intervention
- Develop software solutions and/or improved processes to address work identified as 'toil' by collaborating with key partners to identify, track and remediate processes to free time allocated to reliability
- Partners with Development and Infrastructure teams to create error budget policies prioritizing reliability stories that fall below Service Level Objective (SLO) thresholds and suggests code optimizations, additional instrumentation and/or logging structures to gain service reliability visibility
- Identifies and plans for capacity bottlenecks, vulnerabilities and opportunities for reliability improvement, such as low level error rates and 'noise', and reduces manual support effort and/or improves system reliability
- Assess monitoring for new changes with development partners and works with monitoring tools team to monitor dashboards and enhance application and system monitoring designs
- Engages as a subject matter expert in incident triage efforts, failure scenario modelling and works with the Problem Manager to diagnose root causes for complex/high impact incident/problem management investigations
- Collaborates with Development and Infrastructure teams to understand technical solutions and develop Service Level Indicators and SLOs to measure/improve the reliability of the services they support
- Lead complex platform reliability initiatives such as secondary‑region readiness, egress/ingress observability, private DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation
- Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting for Azure platform services
- Develop reusable Terraform modules, automation frameworks, and CI/CD patterns that improve consistency, compliance, and operational quality
- Drive observability improvements using Azure Monitor, Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools
- Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, automation opportunities, and operational controls
- Partner with security and governance teams to integrate IAM, policy‑as‑code, vulnerability remediation, control validation, and audit readiness into Azure platform operations
- Provide technical design input for new Azure services and workloads to ensure operational readiness before production adoption
- Mentor SRE engineers and raise the technical bar for automation, troubleshooting, documentation, resiliency design, and production support
- Create executive‑ready technical summaries, reliability narratives, and recommendations for leadership review
- Advanced experience in Azure platform engineering, SRE, cloud infrastructure, or enterprise cloud operations
- Deep knowledge of Microsoft Azure architecture, including networking, identity, compute, PaaS, monitoring,…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×