Principal Engineer - Major Incident Response & ITIL Platform
Listed on 2026-09-09
-
IT/Tech
IT Support, IT Project Manager
The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within Service Now and adjacent tooling platforms.
Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.
Key Responsibilities:Major Incident Response (MIR) Program Leadership
- Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
- Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
- Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
- Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
- Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within Service Now.
- Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
- Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within Service Now's CMDB and Problem Management modules.
- Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
- Configure and manage PIR workflows, SLA timers, and notification rules natively in Service Now — no manual handoffs.
- Act as a hands-on technical owner of Service Now ITSM modules:
Incident, Problem, Change, and Event Management. - Design and build Service Now workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, Pager Duty/Alert Ops, etc.).
- Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within Service Now Performance Analytics.
- Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
- Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
- Ensure complete, accurate, and timely documentation of all Problems in Service Now with appropriate categorization and linkage to incidents and changes.
- Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
- Drive a culture of automation-first thinking: identify manual toil and eliminate it through Service Now scripting, Flow Designer, and third-party integrations.
- Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
- Present program health, metrics, and roadmap updates to senior IT and business leadership.
- Faster stabilization of high-severity events through structured, technically informed incident command.
- Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
- A mature, automated Service Now platform that minimizes manual effort and accelerates response and reporting.
- Predictable, trust-building communication to business stakeholders during and after major incidents.
- Continuous improvement embedded into operational DNA — not a periodic exercise.
Impl…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).