More jobs:
Service Manager, Site Reliability Engineering
Job in
Belfast, County Antrim, BT1, Northern Ireland, UK
Listed on 2026-09-18
Listing for:
Allstate Northern Ireland
Full Time
position Listed on 2026-09-18
Job specializations:
-
IT/Tech
IT Support, IT Project Manager, Systems Engineer
Job Description & How to Apply Below
Your role in the team The Service Manager
- Site Reliability Engineering (SRE) is responsible for ensuring the reliability, availability, observability, and operational excellence of technology services while maintaining strong alignment with business objectives. This role serves as the primary owner of service health, incident and problem management, operational governance, continuous service improvement, and stakeholder engagement. The Service Manager partners closely with engineering, platform, infrastructure, security, and business teams to deliver stable, resilient, and high-performing services that align with business objectives and customer expectations.
The role drives proactive risk management, operational maturity, automation, and service improvements while ensuring adherence to established operational processes and governance standards. The role also provides leadership across incident, problem, change, release, and service management disciplines for a team of Service Analysts to deliver resilient, secure, and customer-focused services that meet organizational goals.
Key responsibilities:
Own the overall health, reliability, availability, and observability of critical business applications and technology services, ensuring alignment with established service commitments and customer expectations. Establish, monitor, and report on Service Level Agreements (SLAs), Service Level Objectives (SLOs), Error Budgets, availability, performance, and operational KPIs to drive service excellence. Lead Major Incident Management activities, coordinating cross-functional teams during outages, ensuring effective communication, rapid At Allstate, great things happen when our people work together to protect families and their belongings from life's uncertainties.
And for more than 90 years, our innovative drive has kept us a step ahead of our customers' evolving needs. From advocating for seat belts, air bags and graduated driving laws, to being an industry leader in pricing sophistication, telematics, and, more recently, device and identity protection.
Your role in the team The Service Manager
- Site Reliability Engineering (SRE) is responsible for ensuring the reliability, availability, observability, and operational excellence of technology services while maintaining strong alignment with business objectives. This role serves as the primary owner of service health, incident and problem management, operational governance, continuous service improvement, and stakeholder engagement. The Service Manager partners closely with engineering, platform, infrastructure, security, and business teams to deliver stable, resilient, and high-performing services that align with business objectives and customer expectations.
The role drives proactive risk management, operational maturity, automation, and service improvements while ensuring adherence to established operational processes and governance standards. The role also provides leadership across incident, problem, change, release, and service management disciplines for a team of Service Analysts to deliver resilient, secure, and customer-focused services that meet organizational goals.
Key responsibilities:
Own the overall health, reliability, availability, and observability of critical business applications and technology services, ensuring alignment with established service commitments and customer expectations. Establish, monitor, and report on Service Level Agreements (SLAs), Service Level Objectives (SLOs), Error Budgets, availability, performance, and operational KPIs to drive service excellence. Lead Major Incident Management activities, coordinating cross-functional teams during outages, ensuring effective communication, rapid service restoration, and completion of Root Cause Analysis (RCA) for critical incidents.
Drive Problem Management practices by identifying recurring issues, analyzing systemic failures, implementing permanent corrective actions, and reducing operational risk through preventive measures. Govern Change and Release Management processes by…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×