×
Register Here to Apply for Jobs or Post Jobs. X

Principal Engineer - Major Incident Response & ITIL Platform

Job in Marlborough, Middlesex County, Massachusetts, 01752, USA
Listing for: BJ's Wholesale Club
Full Time position
Listed on 2026-09-09
Job specializations:
  • IT/Tech
    IT Support, IT Project Manager
Salary/Wage Range or Industry Benchmark: 138000 - 176000 USD Yearly USD 138000.00 176000.00 YEAR
Job Description & How to Apply Below

The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within Service Now and adjacent tooling platforms.

Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.

Key Responsibilities:

Major Incident Response (MIR) Program Leadership
  • Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
  • Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
  • Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
  • Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
  • Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within Service Now.
Post-Incident Review & Postmortem Excellence
  • Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
  • Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within Service Now's CMDB and Problem Management modules.
  • Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
  • Configure and manage PIR workflows, SLA timers, and notification rules natively in Service Now — no manual handoffs.
ITIL Platform Engineering & Service Now Ownership
  • Act as a hands-on technical owner of Service Now ITSM modules:
    Incident, Problem, Change, and Event Management.
  • Design and build Service Now workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, Pager Duty/Alert Ops, etc.).
  • Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within Service Now Performance Analytics.
  • Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
  • Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
  • Ensure complete, accurate, and timely documentation of all Problems in Service Now with appropriate categorization and linkage to incidents and changes.
  • Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
  • Drive a culture of automation-first thinking: identify manual toil and eliminate it through Service Now scripting, Flow Designer, and third-party integrations.
  • Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
  • Present program health, metrics, and roadmap updates to senior IT and business leadership.
Key Outcomes:
  • Faster stabilization of high-severity events through structured, technically informed incident command.
  • Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
  • A mature, automated Service Now platform that minimizes manual effort and accelerates response and reporting.
  • Predictable, trust-building communication to business stakeholders during and after major incidents.
  • Continuous improvement embedded into operational DNA — not a periodic exercise.
KPIs & Success Metrics:

Impl…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary