Site reliability engineer
Job in
Greater London, London, Greater London, W1B, England, UK
Listed on 2026-07-23
Listing for:
Enfint
Full Time
position Listed on 2026-07-23
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Описание
EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.
Задачи- Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
- Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
- Define and monitor KPIs for system reliability, performance, and operational efficiency
- Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques
- Develop robust incident management frameworks and lead major incident response activities for critical systems
- Implement blameless postmortems and deliver systemic improvements across production environments
- Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
- Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
- Drive resilience strategies with highly available architectures and disaster recovery readiness
- Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes
- Strong background in Site Reliability Engineering, Dev Ops, or platform operations in complex, distributed environments
- Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights
- Hands‑on experience with automation, Infrastructure as Code (Terraform or Cloud Formation), and CI/CD practices
- Deep understanding of incident management processes, ITSM standards, and ITIL principles
- Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures
- Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability
- Ability to lead transformation, influence across teams, and foster continuous improvement in culture
- Nice to have:
Experience in financial services or other highly regulated, mission‑critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×