Executive Director - Site Reliability Engineering - Retail Pharmacy
Listed on 2026-07-23
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Executive Director, Site Reliability Engineering
We’re building a world of health around every individual – shaping a more connected, convenient and compassionate health experience. At CVS Health®, the Site Reliability Engineering (SRE) organization drives operational excellence across critical healthcare and retail platforms through innovation, automation, observability, and engineering best practices. The Executive Director, Site Reliability Engineering serves as the strategic leader responsible for the reliability, resilience, and performance of CVS Health’s retail and pharmacy technology ecosystem.
This executive defines and executes a comprehensive reliability strategy, oversees large global engineering teams, and establishes a long‑term vision for observability, automation, and operational excellence across thousands of store locations.
- Define and lead the enterprise‑wide Site Reliability Engineering strategy supporting CVS Health’s retail and pharmacy operations. Align reliability and operational objectives with broader business and technology priorities.
- Develop and execute a multi‑year roadmap focused on observability, automation, platform resiliency, and operational excellence.
- Influence technology investment decisions and architectural direction to advance reliability capabilities across the enterprise.
- Serve as the executive sponsor and trusted advisor for store reliability, partnering with retail, pharmacy, digital, and corporate leadership.
- Establish and govern Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets for mission‑critical store applications and services.
- Drive continuous improvements in availability, performance, scalability, and resilience across distributed technology environments.
- Lead enterprise major incident management activities, ensuring rapid detection, effective escalation, root‑cause analysis, and lasting remediation.
- Foster a culture of accountability, operational excellence, and continuous improvement across engineering organizations.
- Establish comprehensive observability capabilities, including advanced monitoring, logging, metrics collection, event correlation, and distributed tracing.
- Deliver end‑to‑end visibility into the health and performance of retail and pharmacy systems across thousands of locations.
- Drive automation initiatives that reduce manual effort, eliminate operational toil, and enable self‑healing operational capabilities.
- Leverage AI and machine learning technologies to improve monitoring, predictive analytics, operational insights, and incident prevention.
- Champion modern cloud, edge computing, and distributed systems architectures that enhance reliability and scalability.
- Partner with Product Engineering, Architecture, Infrastructure, Security, and Operations teams to embed reliability practices throughout the software development lifecycle.
- Ensure technology operations remain secure, compliant, and aligned with healthcare and retail regulatory requirements.
- Lead, mentor, and inspire large global teams of Site Reliability Engineers and engineering leaders.
- Build a high‑performing, inclusive engineering culture focused on innovation, collaboration, and customer outcomes.
- Develop workforce planning strategies and talent pipelines to attract, retain, and grow world‑class engineering talent.
- Promote leadership development and succession planning across the organization.
- 15+ years of progressive technology leadership experience across infrastructure, operations, software engineering, or related disciplines.
- Significant experience leading Site Reliability Engineering, production operations, platform engineering, or large‑scale reliability programs.
- Proven success managing complex, distributed technology environments, preferably in retail, healthcare, pharmacy, or other multi‑site enterprise operations.
- Deep expertise in incident management, observability, automation, resiliency engineering, and operational excellence methodologies.
- Strong understanding of retail technology ecosystems, including POS systems, pharmacy applications, handheld devices, store servers, and network…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).