More jobs:
Senior AIOps and Incident Management/Site Reliability Engineering
Job in
Fort Mill, York County, South Carolina, 29715, USA
Listed on 2026-07-25
Listing for:
Talent Groups
Full Time
position Listed on 2026-07-25
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer
Job Description & How to Apply Below
Senior AIOps Incident Manager / Site Reliability Engineer (SRE)
We are seeking an experienced Senior AIOps Incident Manager / Site Reliability Engineer (SRE) to drive operational excellence, platform reliability, and intelligent automation across a large-scale enterprise environment.
In this role, you will lead major incident management, improve service reliability, implement observability solutions, and champion AIOps initiatives that reduce operational overhead and enhance customer experience. You will collaborate closely with Infrastructure, Cloud, Dev Ops, NOC, Security, and Application Support teams to build resilient, highly available production platforms.
Key Responsibilities- Lead enterprise-wide Major Incident Management and service restoration activities.
- Act as the senior escalation point for critical production incidents.
- Conduct Root Cause Analysis (RCA) and drive corrective and preventive actions.
- Improve operational metrics including MTTR, MTTD, system availability, and service reliability.
- Implement Site Reliability Engineering (SRE) best practices, including SLAs, SLOs, and operational KPIs.
- Design and implement AIOps solutions for intelligent incident detection, event correlation, automated remediation, and self-healing capabilities.
- Lead enterprise observability initiatives using Dynatrace and other monitoring platforms.
- Enhance monitoring across cloud infrastructure, applications, networks, and end-user experience.
- Develop and improve ITIL-based Incident, Problem, Change, and Event Management processes.
- Leverage Service Now workflows to streamline IT service operations.
- Partner with Cloud, Infrastructure, Dev Ops, Security, NOC, and Application teams to improve operational resilience.
- Mentor engineering and operations teams while promoting an automation-first culture.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline (or equivalent experience).
- 8+ years of experience in IT Operations, Infrastructure Operations, Site Reliability Engineering, Production Support, or Network Operations.
- 5+ years leading enterprise incident management, operational transformation, or reliability engineering initiatives.
- Strong experience managing enterprise production environments and mission-critical applications.
- Site Reliability Engineering (SRE)
- IT Service Management (ITSM)
- ITIL Framework
- Problem, Change, Incident, and Event Management
- Network Operations Center (NOC)
- Infrastructure Operations
- Production/Application Support
- Cloud Platforms (AWS, Azure, or GCP)
- Dev Ops practices and CI/CD environments
- Dynatrace or equivalent observability/monitoring platforms
- Service Now (Incident, Change, Problem Management & Workflow Automation)
- Root Cause Analysis (RCA)
- Enterprise Infrastructure, Networking, and Cloud Architecture
Experience with one or more of the following:
- Enterprise AIOps implementations
- AI-powered operational agents
- Intelligent automation and self-healing platforms
- Predictive analytics and anomaly detection
- Workflow orchestration and enterprise automation tools
- Kubernetes, Docker, Terraform, or Ansible
- Splunk, Datadog, App Dynamics, New Relic, Grafana, or Prometheus
- ITIL Foundation or ITIL Managing Professional
- Certified Site Reliability Engineer (SRE)
- AWS, Microsoft Azure, or Google Cloud Certification
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×