More jobs:
Problem Manager – ITIL, ServiceNow, Root Cause Analysis
Job in
Salt Lake City, Salt Lake County, Utah, 84193, USA
Listed on 2026-08-17
Listing for:
Jobtailor
Full Time
position Listed on 2026-08-17
Job specializations:
-
IT/Tech
IT Specialist, IT Support, IT Infrastructure
Job Description & How to Apply Below
- Facilitate timely 5 Whys Root Cause Analysis sessions for P1 incidents and recurring P2 incidents
- Review system logs, monitoring tools, dashboards, and change records to establish timelines and contributing factors
- Use AI-assisted tools for incident investigations, timeline generation, RCA documentation, and knowledge capture
- Document root causes and corrective actions in Service Now Problem Management records
- Lead the end-to-end Problem Management lifecycle from investigation through closure
- Partner with technical teams to create Remediation Action Plans
- Track remediation, accelerate risks, and hold stakeholders accountable
- Validate solutions address root causes and prevent recurrence
- Measure and report service reliability, operational risk, and incident recurrence improvements
- Identify recurring patterns across infrastructure, cloud, application, and network incidents
- Drive proactive problem identification and long-term service improvements
- Recommend monitoring enhancements, automation opportunities, and architectural improvements
- Facilitate weekly Problem Review meetings
- Maintain executive dashboards and Problem Management reporting
- Deliver executive-ready summaries and status updates
- Produce RCA reports for customers and senior leadership
- Improve governance processes, operational standards, runbooks, and reporting practices
- Champion automation and AI-enabled workflows
- 7+ years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related disciplines
- Strong understanding of ITIL Problem Management, Incident Management, and operational governance
- Experience conducting Root Cause Analysis, 5 Whys investigations, and post-incident reviews
- Hands-on experience with Service Now or comparable ITSM platforms
- Strong analytical, documentation, facilitation, and stakeholder management skills
- Ability to influence cross-functional teams and drive accountability without direct authority
- Experience creating executive-level communications, dashboards, and operational reporting
- Broad technical understanding across enterprise technology environments
- Knowledge of Windows and Linux platforms
- Knowledge of VMware and Hyper-V virtualization
- Understanding of performance troubleshooting for CPU, memory, storage, and I/O
- Knowledge of SAN and NAS environments and storage performance, redundancy, and resiliency concepts
- Knowledge of routing, switching, VLANs, DNS, firewalls, load balancing, packet flow, latency, and packet loss
- Knowledge of AWS and Microsoft Azure, including cloud networking, compute, identity, and storage
- Understanding of APIs, microservices, distributed systems, CI/CD, deployment pipelines, application defects, configuration drift, and dependencies
- ITIL Foundation or higher certification preferred
- Experience supporting large-scale enterprise environments preferred
- Experience with operational analytics, trend analysis, and KPI reporting preferred
- Exposure to automation, AI-enabled workflows, or AIOps preferred
- Experience with executive stakeholders and client-facing incident communications preferred
Demonstrates expertise in ITIL Problem Management and Incident Management, with a strong focus on conducting Root Cause Analysis and driving service reliability improvements. Proficient in utilizing Service Now for documentation and reporting, while effectively communicating with executive stakeholders.
Highest-signal resume keywords- ITIL Problem Management
- Root Cause Analysis
- Service Now
- Operational Reporting
- Cloud Networking
Hard Skills
- 5 Whys Investigation
- Operational Analytics
- Performance Troubleshooting
- API Understanding
- CI/CD
- Virtualization (VMware, Hyper-V)
- Storage Performance (SAN, NAS)
- Networking (Routing, Switching, VLANs)
- Cloud Platforms (AWS, Microsoft Azure)
- Documentation
- Analytical Skills
- Facilitation Skills
- Stakeholder Management
- Influencing Skills
- Communication Skills
- ITIL Foundation
- IT Operations
- Incident Management
- Major Incident Management
- Operational Governance
- Service Reliability
- Service Now
- Monitoring Tools
- Dashboards
- AI-Assisted Tools
- Executive Dashboards
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×