SME Shift Supervisor
Listed on 2026-09-30
-
IT/Tech
Systems Administrator, IT Support, Cloud Computing: Infrastructure & Operations, Technical Support
Position Overview
The SME Shift Supervisor serves as part of the incident management team in a cloud-based environment, diagnosing, mitigating, and/or escalating system issues to maintain a high level of system/platform availability. This role requires an understanding of core cloud system components and tools to diagnose issues, and acts as an escalation point for more complex incidents. The SME Shift Supervisor responds to incident tickets to meet SLA objectives, typically handling the more complex incidents, and supports collaboration across operations, development teams, and external partners to drive timely resolution and continuous improvement.
Key ResponsibilitiesServe as part of the incident management team in a cloud-based environment, diagnosing, mitigating, and/or escalating system issues to maintain high system/platform availability.
Act as an escalation point for more complex incidents.
Respond to incident tickets in an operational environment to meet SLA objectives, typically handling the more complex incidents.
Troubleshoot system issues using diagnostic tools such as Netmon, Win Dbg, and custom application tools.
Review system logs to identify and mitigate system issues.
Leverage the knowledge base to help troubleshoot, identify, and resolve systems issues.
Update knowledge base troubleshooting guides and lessons learned as required.
Document incident fixes and make recommendations to the engineering team for system improvements for consideration in future releases.
Document system issues resulting in system outages and coordinate change through the change management process.
Support collaboration across operations, development teams, and external partners.
Support "Tiger team" calls to streamline knowledge sharing and timely resolution of system issues.
Monitor solution performance according to client specifications and SLAs.
Serve as an escalation point on more complex issues.
BS in Computer Science or other technical discipline preferred.
3 years of operations experience providing application infrastructure support; 2 years performing system administrator support.
Experience with system administration support tools such as Windows/Linux.
Experience supporting a cloud-based environment, including tools such as Azure/AWS.
Experience analyzing, troubleshooting, and providing solutions for technical issues.
Strong interpersonal skills; strong oral and written communication skills.
Strong technical communication with both technical and non-technical peers.
Ability to problem-solve and collaborate with team members.
Strong organizational and multi-tasking skills.
Able to maintain professionalism under pressure; strong customer focus.
Ability to work onsite and support the on-call weekend supervisor rotation.
Top-Secret Clearance with SCI and Full Scope Polygraph.
Active Passport/Passbook.
No dual citizen ships.
Relevant cloud/systems certifications (e.g., Azure/AWS certifications, Security+, Network+, ITIL/Problem Management).
Prior experience in a customer-facing government operational environment.
Prior experience supporting incident management within a formal SLA-driven operations environment.
Incident diagnosis, mitigation, and escalation in a cloud environment.
Diagnostic tooling:
Netmon, Win Dbg, custom application tools.System log analysis and root-cause troubleshooting.
Knowledge base management (troubleshooting guides, lessons learned).
Change management documentation and coordination.
Cross-team collaboration (operations, development, external partners).
SLA/solution performance monitoring.
"Tiger team" knowledge-sharing facilitation.
- Day shift, Monday-Friday, 8-hour shifts; supports the on-call weekend supervisor…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).