Principal Site Reliability Engineer
Listed on 2026-07-17
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Disaster Recovery IT, Systems Engineer
The Platform Engineering team builds, secures and operates scalable infrastructure supporting cloud-managed SaaS products with on-premises components deployed at customer sites.
As our Principal Site Reliability Engineer, you will provide the technical leadership for reliability across our global platform. You will define the strategy, standards and operating model that ensure highly available, resilient and secure services, while remaining hands‑on with the design and operation of the technologies that underpin our production environments.
Working closely with Architecture, Dev Sec Ops , Cloud Operations and Product Development, you will drive a culture of observability, automation and continuous improvement, ensuring issues are identified and resolved before they impact customers.
- Define and lead the reliability strategy for Tier 1 and Tier 2 production services, including service-level objectives, error budgets, observability, resilience, disaster recovery and cloud security.
- Design and operate observability platforms, lead chaos engineering and disaster recovery initiatives, improve the reliability of PostgreSQL, Redis/Valkey, Kafka and Open Search, and deliver AI Ops capabilities including predictive monitoring, automated remediation and self‑healing.
- Lead major incident response, on‑call operations and global escalation, while mentoring engineers and establishing engineering standards adopted across the organisation.
You will own the reliability, resilience and operational excellence of our cloud platform and production services, embedding reliability principles into architectural decisions and platform design. You'll establish observability standards covering metrics, logs, distributed tracing and profiling, while governing service-level objectives and error budgets across engineering teams.
You will lead resilience engineering through chaos testing, disaster recovery planning and validated failover exercises, co‑own cloud security posture with Dev Sec Ops , and ensure the reliability of our critical data and streaming platforms. Working across multiple engineering disciplines, you will also optimise platform efficiency through automation, AI‑driven operations and continuous operational improvement.
- Own production reliability, observability, service-level objectives, error budgets, resilience, backup and disaster recovery, cloud security and platform governance across all critical services.
- Lead incident command, executive communications, post‑incident reviews, on‑call operations and global escalation, while delivering predictive monitoring, automated remediation and self‑healing capabilities.
- Partner with Architecture, Dev Sec Ops , Cloud Operations and Product Development to establish engineering standards, mentor engineers and deliver a shared reliability roadmap.
You are an experienced Site Reliability Engineering leader with more than 15 years of production engineering experience and recent hands‑on responsibility for large‑scale, fault‑tolerant production systems running on AWS or GCP.
You have successfully designed and operated observability platforms, implemented service-level objective programmes and delivered measurable improvements in availability, reliability and mean time to recovery. You have led high‑severity production incidents, planned and executed resilience testing and disaster recovery exercises, and influenced engineering practices across multiple teams through technical leadership and mentoring.
- 15+ years of production engineering experience with recent hands‑on responsibility for large‑scale, fault‑tolerant production systems on AWS or GCP, including ownership of observability, service-level objectives and error budgets.
- Proven experience leading resilience initiatives, validated disaster recovery and failover testing, together with command of high‑severity production incidents and measurable reliability improvements.
- Demonstrated technical leadership through architecture reviews, governance, mentoring, coaching and engineering standards adopted across multiple teams and services.
You have deep expertise across…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: