×
Register Here to Apply for Jobs or Post Jobs. X

Site reliability engineer

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Enfint
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 180000 GBP Yearly GBP 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Описание

EPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.

Задачи
  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes
Требования
  • Strong background in Site Reliability Engineering, Dev Ops, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or Cloud Formation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures
  • Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture
  • Nice to have:
    Experience in financial services or other highly regulated, mission‑critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary