×
Register Here to Apply for Jobs or Post Jobs. X

Senior AIOps and Incident Management​/Site Reliability Engineering

Job in Fort Mill, York County, South Carolina, 29715, USA
Listing for: Talent Groups
Full Time position
Listed on 2026-07-25
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 180000 USD Yearly USD 140000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Senior AIOps and Incident Management /Site Reliability Engineering

Senior AIOps Incident Manager / Site Reliability Engineer (SRE)

We are seeking an experienced Senior AIOps Incident Manager / Site Reliability Engineer (SRE) to drive operational excellence, platform reliability, and intelligent automation across a large-scale enterprise environment.

In this role, you will lead major incident management, improve service reliability, implement observability solutions, and champion AIOps initiatives that reduce operational overhead and enhance customer experience. You will collaborate closely with Infrastructure, Cloud, Dev Ops, NOC, Security, and Application Support teams to build resilient, highly available production platforms.

Key Responsibilities
  • Lead enterprise-wide Major Incident Management and service restoration activities.
  • Act as the senior escalation point for critical production incidents.
  • Conduct Root Cause Analysis (RCA) and drive corrective and preventive actions.
  • Improve operational metrics including MTTR, MTTD, system availability, and service reliability.
  • Implement Site Reliability Engineering (SRE) best practices, including SLAs, SLOs, and operational KPIs.
  • Design and implement AIOps solutions for intelligent incident detection, event correlation, automated remediation, and self-healing capabilities.
  • Lead enterprise observability initiatives using Dynatrace and other monitoring platforms.
  • Enhance monitoring across cloud infrastructure, applications, networks, and end-user experience.
  • Develop and improve ITIL-based Incident, Problem, Change, and Event Management processes.
  • Leverage Service Now workflows to streamline IT service operations.
  • Partner with Cloud, Infrastructure, Dev Ops, Security, NOC, and Application teams to improve operational resilience.
  • Mentor engineering and operations teams while promoting an automation-first culture.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline (or equivalent experience).
  • 8+ years of experience in IT Operations, Infrastructure Operations, Site Reliability Engineering, Production Support, or Network Operations.
  • 5+ years leading enterprise incident management, operational transformation, or reliability engineering initiatives.
  • Strong experience managing enterprise production environments and mission-critical applications.
Required Technical Skills
  • Site Reliability Engineering (SRE)
  • IT Service Management (ITSM)
  • ITIL Framework
  • Problem, Change, Incident, and Event Management
  • Network Operations Center (NOC)
  • Infrastructure Operations
  • Production/Application Support
  • Cloud Platforms (AWS, Azure, or GCP)
  • Dev Ops practices and CI/CD environments
  • Dynatrace or equivalent observability/monitoring platforms
  • Service Now (Incident, Change, Problem Management & Workflow Automation)
  • Root Cause Analysis (RCA)
  • Enterprise Infrastructure, Networking, and Cloud Architecture
Preferred Qualifications

Experience with one or more of the following:

  • Enterprise AIOps implementations
  • AI-powered operational agents
  • Intelligent automation and self-healing platforms
  • Predictive analytics and anomaly detection
  • Workflow orchestration and enterprise automation tools
  • Kubernetes, Docker, Terraform, or Ansible
  • Splunk, Datadog, App Dynamics, New Relic, Grafana, or Prometheus
Preferred Certifications
  • ITIL Foundation or ITIL Managing Professional
  • Certified Site Reliability Engineer (SRE)
  • AWS, Microsoft Azure, or Google Cloud Certification
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary