×
Register Here to Apply for Jobs or Post Jobs. X

Senior Manager, Site Reliability & Operational Resilience

Job in Morristown, Morris County, New Jersey, 07960, USA
Listing for: LE018 Zelis Healthcare, LLC
Full Time position
Listed on 2026-08-17
Job specializations:
  • IT/Tech
    IT Project Manager
  • Management
    IT Project Manager
Salary/Wage Range or Industry Benchmark: 139000 - 176700 USD Yearly USD 139000.00 176700.00 YEAR
Job Description & How to Apply Below

At Zelis, we Get Stuff Done. So, lets get to it!

A Little About Us

Zelis is modernizing the healthcare financial experience across payers, providers, and healthcare consumers. We serve more than 750 payers, including the top five national health plans, regional health plans, TPAs and millions of healthcare providers and consumers across our platform of solutions. Zelis sees across the system to identify, optimize, and solve problems holistically with technology built by healthcare experts - driving real, measurable results for clients.

A

Little About You

You bring a unique blend of personality and professional expertise to your work, inspiring others with your passion and dedication. Your career is a testament to your diverse experiences, community involvement, and the valuable lessons you've learned along the way. You are more than just your resume; you are a reflection of your achievements, the knowledge you've gained, and the personal interests that shape who you are.

Position Overview

The Senior Manager, Site Reliability & Operational Resilience will lead and mature the enterprise capabilities that enable Zelis to detect, respond to, recover from, and continuously learn from technology disruptions. Reporting to the Director, Global Operations, this leader will set the strategy and operating model for Enterprise Observability, Major Incident Command, Disaster Recovery Orchestration, and reliability engineering, while partnering with the IT Service Management Process team to strengthen problem management and drive disciplined execution of major-incident corrective actions.

This leader will lead a globally distributed function in close partnership with an India-based leader. Together, they will align priorities, standards, coverage, handoffs, performance measures, and talent development as one global organization. This is a build-and-transform opportunity for a technically credible, pragmatic leader who enjoys fixing what is not working, creating durable operating mechanisms, and scaling strong practices across a complex enterprise.

The successful candidate will combine calm leadership under pressure with the engineering depth, influence, and persistence required to turn reliability and resilience into measurable business outcomes.

What Youll Do

Build and scale the practice.

Define and execute a multi-year Site Reliability & Operational Resilience roadmap, including the target operating model, service offerings, governance, standards, talent plan, maturity measures, and adoption strategy required to operate at enterprise scale.

Lead a global team of senior engineers.

Coach, organize, and develop a team composed primarily of senior engineers and technical leads.

Partner with the India-based leader to create clear ownership, effective follow-the-sun handoffs, sustainable coverage, strong technical decision-making, career growth, and a culture of high autonomy with clear accountability.

Own the enterprise observability strategy.

Establish the target-state architecture and operating model across Logic Monitor, New Relic, Splunk, and Datadog. Standardize telemetry across metrics, logs, traces, events, synthetic monitoring, and service health; improve onboarding, dashboards, integration, signal quality, alert precision, platform economics, and adoption across critical services.

Mature the Major Incident Command capability.

Lead, coach, and scale the Incident Commander function. Establish a consistent command model, severity standards, decision rights, playbooks, technical and business coordination, global handoffs, executive communications, and learning mechanisms that accelerate service restoration and increase confidence during high-impact events.

Build Disaster Recovery Orchestration.

Create the process, governance, annual testing strategy, roles, communications, and cross-functional coordination needed to execute reliable disaster recovery exercises. Establish and govern a single source of truth for recovery plans, runbooks, dependencies, ownership, test evidence, lessons learned, and remediation status. Make recovery readiness visible and measurable.

Define recovery-readiness measures and dashboards that…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary