Senior Manager, Site Reliability & Operational Resilience
Listed on 2026-08-16
-
IT/Tech
Disaster Recovery IT, SRE/Site Reliability
Senior Manager, Site Reliability & Operational Resilience
The Senior Manager, Site Reliability & Operational Resilience will lead and mature the enterprise capabilities that enable Zelis to detect, respond to, recover from, and continuously learn from technology disruptions. Reporting to the Director, Global Operations, this leader will set the strategy and operating model for Enterprise Observability, Major Incident Command, Disaster Recovery Orchestration, and reliability engineering, while partnering with the IT Service Management Process team to strengthen problem management and drive disciplined execution of major-incident corrective actions.
This leader will lead a globally distributed function in close partnership with an India-based leader. Together, they will align priorities, standards, coverage, handoffs, performance measures, and talent development as one global organization. This is a build-and-transform opportunity for a technically credible, pragmatic leader who enjoys fixing what is not working, creating durable operating mechanisms, and scaling strong practices across a complex enterprise.
The successful candidate will combine calm leadership under pressure with the engineering depth, influence, and persistence required to turn reliability and resilience into measurable business outcomes.
What You'll Do
- Build and scale the practice. Define and execute a multi-year Site Reliability & Operational Resilience roadmap, including the target operating model, service offerings, governance, standards, talent plan, maturity measures, and adoption strategy required to operate at enterprise scale.
- Lead a global team of senior engineers. Coach, organize, and develop a team composed primarily of senior engineers and technical leads. Partner with the India-based leader to create clear ownership, effective follow-the-sun handoffs, sustainable coverage, strong technical decision-making, career growth, and a culture of high autonomy with clear accountability.
- Own the enterprise observability strategy. Establish the target-state architecture and operating model across Logic Monitor, New Relic, Splunk, and Datadog. Standardize telemetry across metrics, logs, traces, events, synthetic monitoring, and service health; improve onboarding, dashboards, integration, signal quality, alert precision, platform economics, and adoption across critical services.
- Mature the Major Incident Command capability. Lead, coach, and scale the Incident Commander function. Establish a consistent command model, severity standards, decision rights, playbooks, technical and business coordination, global handoffs, executive communications, and learning mechanisms that accelerate service restoration and increase confidence during high-impact events.
- Build Disaster Recovery Orchestration. Create the process, governance, annual testing strategy, roles, communications, and cross-functional coordination needed to execute reliable disaster recovery exercises. Establish and govern a single source of truth for recovery plans, runbooks, dependencies, ownership, test evidence, lessons learned, and remediation status.
- Make recovery readiness visible and measurable. Define recovery-readiness measures and dashboards that monitor plan currency, test coverage, critical dependencies, RTO/RPO attainment, recovery gaps, and remediation aging. Help advance the organization from periodic disaster recovery testing toward continuous operational resilience through scenario exercises, game days, failover validation, and ongoing learning.
- Close the loop after incidents. Partner with the IT Service Management Process team to improve post-incident reviews, root-cause quality, known-error practices, and execution of problem-management and major-incident action items. Create transparent mechanisms for ownership, due dates, dependencies, aging, risk acceptance, escalation, and verification of effectiveness while keeping delivery accountability with the assigned action owners.
- Establish reliability and resilience standards. Define practical standards for service tiering, SLIs, SLOs, error budgets, production readiness, capacity, dependency management, recovery objectives, resilience testing, operational health, and reliability reviews. Embed these practices into the lifecycle of critical services.
- Engineer out toil and recurring failure. Turn operational pain points into an engineering backlog and drive automation, runbook automation, self-service, event correlation, self-healing, and prioritized technical-debt remediation that reduce manual work and prevent repeat incidents.
- Influence across the enterprise. Partner with Application Engineering, Infrastructure, Cloud and Platform Engineering, Cybersecurity, Enterprise Architecture, Business Continuity, Risk and Compliance, Product, business operations, and third-party providers to embed reliability, recoverability, and resilience into technology decisions and service ownership.
- Measure and communicate what matters. Create…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).