Director, Site Reliability Engineering - Incident Management
Listed on 2026-09-05
-
IT/Tech
Business Continuity
Director, Site Reliability Engineering – Incident Management
The Director, Site Reliability Engineering – Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship.
This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.
~30+ Teams – Byte, KFC and Taco Bell
Global Team - Vietnam, India, Colombia and US
Platform Team – Responsible for all products (Edge, Commerce, POS, KDS, Menu, Portal, etc.)
ResponsibilitiesPosition Functions:
Strategic Leadership
- Provide strategic leadership for Incident Management and Technical Operations across Byte, KFC, and Taco Bell digital platforms, ensuring high availability, operational excellence, and a consistent customer experience.
- Establish the vision, governance, and operating model for enterprise Incident Management, driving standardized processes, tooling, and best practices across multiple brands and technology organizations.
- Drive enterprise operational readiness for major product launches, restaurant initiatives, promotions, and seasonal events through effective change management, risk assessment, and cross-functional planning.
- Own the enterprise framework and targets for service level objectives and indicators, set in partnership with the engineering teams that own the services, which remain accountable for achieving them. Define and monitor MTTR, incident trends, service health, operational maturity, and executive KPIs, using data to prioritize investments and improve platform resilience.
- Own the enterprise Incident Management governance framework, ensuring consistent execution, accountability, and continuous improvement across all brands.
- Own the enterprise observability strategy, setting the standard and direction for how platform health is measured and seen, with Platform Engineering partnering on the underlying platform and instrumentation.
- Own the strategy and roadmap for the reliability platform, including the tooling, products, and self-service capabilities that deliver reliability to engineering teams and to markets.
- Provide executive-level communications and operational updates to Digital & Technology leadership, Brand CDTOs/CTOs, and executive stakeholders during major incidents and through regular operational business reviews.
- Develop trusted partnerships with executive leaders across Product, Engineering, Infrastructure, Security, Restaurant Operations, and Brand Technology to align operational priorities with business objectives and customer experience goals.
- Jointly own Business Continuity and Disaster Recovery with Platform Engineering, covering recovery strategies, resiliency testing, crisis management processes, and operational preparedness across Byte, KFC, and Taco Bell. Decisions are made in a standing joint review, with anything unresolved escalating to the Senior Director, Engineering within one cycle. During a declared continuity event or major incident, the Incident Commander decides in the moment.
- Sponsor continuous improvement initiatives that enhance operational resilience, reduce enterprise risk, and strengthen the organization's ability to respond to large-scale operational events.
Technical…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).