Cloud Operations - Service Reliability Engineer
Job in
Holywood, County Down, BT18, Northern Ireland, UK
Listed on 2026-08-03
Listing for:
A&O Shearman
Full Time
position Listed on 2026-08-03
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure Dev Ops, Git Hub and Ansible to support consistent, repeatable and well-governed service operation.
Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;
Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; and
Promoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
The role involves:
Supporting service reliability, observability and cloud engineering across the following areas:
Monitoring, observability and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals and operational events – Essential;
Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC and platform diagnostics – Essential;
Infrastructure as Code and automation using Bicep, Azure Dev Ops pipelines and Git Hub-based source control and collaboration – Essential;
Configuration management and standards automation using Ansible or equivalent tooling – Preferred;
Experience of using or implementing monitoring solutions using Elastic – Preferred;
Operational reporting, issue trend analysis and the development of actionable dashboards to support service improvement – Preferred;
Experience of working across a broad range of systems, technologies and internal support teams – Preferred.
Ensuring that monitoring and operational insight are effectively designed, implemented and understood so that services can be supported, improved and made more resilient.
Providing subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
Working globally across cloud-hosted services and platform capabilities, independent of location.
Support the firm’s environmental goals and initiatives.
Monitoring, Reliability and Cloud Engineering
Works with internal technology teams to improve end-to-end observability for supported services, including:
Monitoring coverage for infrastructure, platform services and application components;
Actionable alerting that supports early identification of degradation, failure or operational risk;
Dashboards and reporting that help teams understand service health, trends and recurring issues;
Cloud engineering practices that use Bicep, Azure Dev Ops, Git Hub and Ansible to deliver consistent and repeatable change; and
Operational standards that improve service resilience and reduce manual support effort.
Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.
Service Delivery
Identify, diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
Work with service-owning teams to improve the quality, relevance and routing of alerts so that operational issues can be detected and acted upon quickly.
Contribute to root cause analysis, problem management and continuous improvement activity by identifying recurring patterns, gaps in observability and opportunities for automation.
Build and Implementation
Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
Support implementation of monitoring and automation standards across new and existing services.
Ensure that operational documentation, handover materials and support guidance are created and are suitable for BAU operation.
Risk Management
Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
Refer to domain experts for guidance on specialised areas such as architecture, security, networking,…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×