Incident Manager
Job in
Philadelphia, Philadelphia County, Pennsylvania, 19117, USA
Listed on 2026-07-27
Listing for:
Aceolution
Full Time
position Listed on 2026-07-27
Job specializations:
-
IT/Tech
SRE/Site Reliability, IT Support, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
We are seeking an experienced Incident Manager III to lead critical incident response, problem management, and operational reliability initiatives. The ideal candidate will have strong technical expertise in cloud technologies, infrastructure, monitoring tools, and incident management, along with the ability to coordinate cross-functional teams during high-severity production incidents.
Key Responsibilities
- Lead high-severity P0/P1 incidents as the Incident Commander.
- Coordinate cross-functional engineering teams during critical production incidents.
- Drive rapid troubleshooting, impact assessment, and resolution decisions.
- Prepare executive-ready incident summaries and ensure accurate documentation.
- Apply technical understanding of application architecture and service dependencies to accelerate incident mitigation.
- Lead investigations for recurring and major incidents.
- Perform detailed Root Cause Analysis (RCA) with engineering teams.
- Validate corrective actions and monitor long-term remediation plans.
- Present findings, lessons learned, and preventive strategies to leadership.
Change Management
- Review and approve high-risk or complex production changes.
- Participate in Change Advisory Board (CAB) meetings as required.
- Validate rollback strategies and operational readiness.
- Lead change execution for major maintenance activities and reliability events.
Reliability & Operational Excellence
- Drive initiatives to improve MTTA, MTTR, and overall service reliability.
- Mentor junior team members on incident response best practices.
- Partner with SRE and Engineering teams to enhance observability, automation, and resilience.
- Lead maintenance events, disaster recovery exercises, failover testing, and resilience validation.
- Review and improve runbooks, monitoring strategies, and automation workflows.
- Contribute ideas around AI and automation to improve operational efficiency and reduce MTTM/MTTR.
Required Qualifications
- Bachelor's degree in Engineering, Computer Science, or a related technical discipline.
- 3–5 years of experience in Site Reliability Engineering (SRE), Incident Management, Operations, Dev Ops, or a similar role.
- Strong expertise in monitoring tools, logging platforms, and incident response processes.
- Excellent troubleshooting skills across networking, Linux/Windows servers, cloud infrastructure, and distributed systems.
- Hands-on experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Experience with automation using Python, Go, Bash, or Power Shell.
- Good understanding of microservices, containers, and distributed architectures.
- Excellent communication, leadership, and stakeholder management skills with the ability to perform under pressure.
Preferred Skills
- Experience working in enterprise production support environments.
- Knowledge of ITIL Incident, Problem, and Change Management processes.
- Exposure to observability platforms and automation frameworks.
- Experience supporting large-scale cloud-native applications.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×