Senior DevOps & Site Reliability Engineer
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
We are looking for an experienced
Senior Dev Ops & Site Reliability Engineer (SRE)to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.
This is a senior hands-on engineering role spanning
Dev Ops, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and Dev Sec Ops .
The successful candidate will work across engineering and delivery teams to improve
platform reliability, deployment velocity, resilience, automation, operational efficiency and production performance
, while supporting mission-critical enterprise applications.
Design, build and maintain
cloud-native infrastructure and platform services
.Develop and maintain
Infrastructure as Code (IaC)solutions.Automate infrastructure provisioning, configuration and operational processes.
Build reusable engineering tools, deployment templates and platform components.
Establish and standardise platform engineering practices across multiple delivery teams.
Identify opportunities to reduce manual intervention and increase engineering automation.
Design, implement and maintain enterprise-grade
CI/CD pipelines
for application and infrastructure deployments.Implement automated testing, security scanning, code-quality controls and release automation.
Enable automated deployments, rollback and recovery processes.
Improve deployment frequency while reducing change and deployment risk.
Continuously optimise software delivery and release-management processes.
Implement and mature
Site Reliability Engineering practices
across production environments.Define, monitor and manage
Service Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs).Improve application and platform
availability, scalability, resilience and performance
.Lead production incident response, troubleshooting, problem management and
Root Cause Analysis (RCA).Drive proactive reliability improvements and reduction of technical debt.
Improve
Mean Time to Detect (MTTD)and
Mean Time to Recover (MTTR).
Design, implement and operate enterprise
Microsoft Azure
environments.Work extensively with technologies such as:
Azure Kubernetes Service (
AKS
)Azure App Services
Azure Networking
Azure Monitor
Azure Storage
Azure Identity Services
Design highly available and disaster-recovery-capable environments.
Optimise cloud environments for
performance, resilience, security and cost
.Support hybrid-cloud and multi-cloud environments where required.
Build, deploy and support containerised applications using
Docker
and
Kubernetes
.Manage Kubernetes environments, particularly
Azure Kubernetes Service (AKS).Develop and maintain deployment configurations using
Helm
.Support container-platform reliability, scalability and operational performance.
Open Shift experience would be advantageous.
Hands-on experience with technologies such as:
Terraform
Bicep
ARM Templates
Ansible
Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.
Monitoring & ObservabilityImplement comprehensive
logging, monitoring, metrics, tracing and alerting
.Build operational dashboards and platform insights.
Establish enterprise observability standards.
Implement proactive and predictive monitoring.
Use observability information to improve application and infrastructure reliability.
Relevant technologies may include:
Dynatrace
Grafana
Prometheus
Elastic Stack / ELK
Splunk
Azure Monitor
Open Telemetry
Embed
Dev Sec Ops practices throughout the software-delivery lifecycle.Integrate security scanning and controls into CI/CD pipelines.
Support vulnerability identification, remediation and risk reduction.
Ensure cloud and platform environments comply with enterprise security and regulatory requirements.
Work closely with information-security teams to continuously improve platform security.
Provide technical leadership across Dev Ops, Cloud, Platform and SRE…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: