Incident Manager
Listed on 2026-08-29
-
IT/Tech
IT Support, Systems Administrator, Systems Analyst, IT Project Manager
Incident Manager
Koniag IT Systems, LLC a Koniag Government Services company, is seeking an Incident Manager to support KITS and our government customer Washington, DC. The position is hybrid, requires 3 days onsite. This position requires the candidate to be able to obtain a Public Trust. We offer competitive compensation and an extraordinary benefits package including health, dental and vision insurance, 401K with company matching, flexible spending accounts, paid holidays, three weeks paid time off, and more.
The Incident Manager operates within a mature ITIL 4-aligned IT Service Management (ITSM) practice and works in close coordination with the enterprise NOC, multi-tier Service Desk, network engineering, infrastructure, and platform operations teams to ensure that all incidents are managed to resolution within established service level objectives. This individual plays a central role in driving service restoration, minimizing business impact, and contributing to the program's continual service improvement objectives through trend analysis, problem identification, and post-incident learning activities.
The ideal candidate is a detail-oriented and operationally experienced ITSM professional with a strong background in enterprise incident management, ITIL-based service management practices, and federal IT operations environments. This individual must thrive in a fast-paced, high-accountability 24x7 operational setting and possess the ability to coordinate effectively across diverse technical teams and Government stakeholders under time-critical conditions.
Principal responsibilities will include but are not limited to:
- Own and manage the full enterprise incident management lifecycle from initial detection and logging through triage, categorization, prioritization, assignment, escalation, resolution, and formal closure in the Government's ITSM platform.
- Ensure all incidents are accurately classified by priority (P1-P4) in accordance with established impact and urgency criteria, and that appropriate response and resolution timelines are applied and monitored in real time.
- Monitor the active incident queue continuously, identifying aging incidents, escalation triggers, SLA breach risks, and stalled assignments, and taking proactive action to ensure forward progress on all open incidents.
- Coordinate and drive the timely assignment of incidents to the appropriate Tier 1, Tier 2, Tier 3, or Tier 4 technical resources based on incident category, complexity, and required expertise.
- Ensure incidents are escalated promptly and appropriately-both functionally across technical teams and hierarchically to program leadership and Government stakeholders-when SLA thresholds, scope, or business impact warrant escalation.
- Maintain accurate, complete, and timely incident records in the ITSM platform throughout the incident lifecycle, ensuring all actions, decisions, workarounds, and resolutions are thoroughly documented.
Major Incident Management:
- Serve as the primary coordinator and decision-maker for all Priority 1 (P1) and Priority 2 (P2) major incidents, activating and leading technical bridge calls, war rooms, and cross-functional response teams to drive rapid service restoration.
- Ensure P1 critical incidents are acknowledged within 15 minutes and that service restoration is achieved within 4 hours in accordance with program service level objectives, meeting the Government's target of at least 90% of P1 incidents restored within the 4-hour SLA.
- Maintain clear, accurate, and timely stakeholder communications throughout major incident events, including initial notifications, status updates at defined intervals, and formal incident closure notifications to the COR and Government stakeholders.
- Coordinate with the NOC, Service Desk Manager, technical leads, and program leadership to ensure all required resources are engaged, barriers to resolution are removed, and escalation paths are followed during major incidents.
- Lead or facilitate formal post-incident reviews (PIRs) for all P1 and significant P2 incidents, ensuring root cause analysis findings, contributing factors, corrective actions, and lessons learned are thoroughly documented and tracked to completion.
Problem Management Coordination:
- Collaborate with the Problem Manager or Service Management Lead to ensure incidents exhibiting patterns of recurrence or systemic root causes are formally transitioned into the problem management process.
- Contribute incident trend data, recurring issue patterns, and workaround documentation to support problem record creation, root cause investigation, and known error documentation.
- Track and monitor the implementation of permanent fixes and corrective actions arising from problem management activities, ensuring incident recurrence rates are measurably reduced over time.
Change Management Support:
- Coordinate with the change management process to identify incidents resulting from unauthorized changes or failed changes, ensuring accurate linkage between…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).