Site Reliability Engineer
Listed on 2026-09-06
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Disaster Recovery IT, Azure
Willis Re is expanding its Global business in 2026 and the Cloud and Infrastructure team will need to support this growth by designing provisioning then supporting the platforms to enable this.
We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing operating automating and continuously improving enterprise-scale Azure platforms ensuring high availability resiliency security and performance.
The role combines software engineering cloud architecture infrastructure automation and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices.
The ideal candidate will have strong experience with Microsoft Azure
Dev Ops
Terraform
API disaster recovery planning and enterprise-scale resilience engineering
.
Key Responsibilities
Platform Reliability & Operations- Ensure the availability performance scalability and reliability of Azure-hosted services.
- Define and manage Service Level Indicators (SLIs) Service Level Objectives (SLOs) and error budgets.
- Proactively monitor platform health and performance using observability tooling.
- Perform root cause analysis and implement permanent fixes for recurring incidents.
- Participate in incident management and on-call support rotations where required.
- Lead blameless post-incident reviews capture lessons learned and drive corrective actions through to completion.
- Reduce operational toil by identifying repetitive manual tasks and replacing them with automated reusable engineering solutions.
- Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.
- Design deploy and manage Azure infrastructure services including:
- Virtual Networks
- Application Gateways
- API Management
- Azure Kubernetes Service (AKS)
- Azure Firewall
- Azure Storage
- Key Vault
- Azure Monitor
- Azure AI Services
- Implement cloud platform standards and best practices.
- Support multi-region Azure deployments and platform modernisation initiatives.
- Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.
- Develop and maintain Terraform modules and reusable infrastructure patterns.
- Implement Infrastructure as Code (IaC) standards and governance controls.
- Ensure infrastructure is version controlled peer-reviewed and fully automated.
- Manage Terraform state securely and consistently across environments.
- Build and maintain Azure Dev Ops CI/CD pipelines.
- Automate infrastructure provisioning and application deployments.
- Implement testing security scanning policy compliance and release gates.
- Support Git Ops and platform engineering practices.
- Create automation for operational runbooks self-healing processes deployment validation and environment consistency checks.
- Design and implement highly available Azure architectures.
- Develop and maintain disaster recovery and business continuity capabilities.
- Implement and test:
- Regional failover strategies
- Active/Passive architectures
- Active/Active deployments
- Traffic Manager and Front Door failover patterns
- Database resiliency and replication
- Backup and recovery solutions
- Conduct regular resilience and recovery testing exercises.
- Identify and reduce single points of failure across platforms.
- Define and execute game days chaos testing and controlled failure scenarios to validate operational resilience.
- Ensure platforms are secure-by-design.
- Work closely with Security and Architecture teams to implement:
- Zero Trust principles
- RBAC controls / Managed Identities
- Network segmentation
- Secrets management
- Support compliance requirements and operational audits.
- Help coordinate security updates patches maintenance routines and upgrades of the underlying system across partners and vendors
- Embed reliability security and compliance controls into build and release pipelines to support production readiness.
- Drive automation and reduction of manual operational…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: