Senior Client-Facing SRE/Cloud Engineer; AWS, Kubernetes
Listed on 2026-09-03
-
IT/Tech
SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations
Location: Greater London
B2B Contract | EU or US - fully remote
Role OverviewWe are looking for a Senior Cloud Reliability Engineer to take hands‑on ownership of highly available, cloud-native production environments.
This is not a traditional Dev Ops role focused primarily on building CI/CD pipelines or migrating infrastructure. We are looking for an engineer who has operated critical production systems, owned incidents while on call, and can identify weaknesses in an existing cloud environment and drive meaningful improvements.
The role combines deep AWS and Kubernetes engineering with SRE practices, infrastructure automation, observability, incident management, and direct technical interaction with enterprise customers.
You will have significant autonomy to challenge existing approaches, propose better solutions, and improve the reliability, scalability, security, and operational maturity of the platform.
Key Responsibilities- Own the reliability and operational health of production AWS and Kubernetes environments.
- Participate in on-call rotations and take ownership of high-severity production incidents from detection through mitigation and resolution.
- Lead root cause analysis and post-incident reviews, implementing permanent corrective actions rather than temporary fixes.
- Identify architectural, reliability, security, performance, and operational weaknesses within existing cloud environments.
- Propose and implement improvements based on AWS and Kubernetes best practices.
- Design, maintain, and continuously improve AWS infrastructure and production Kubernetes/EKS environments.
- Automate infrastructure provisioning and operational workflows using Terraform and configuration-management tools.
- Improve deployment and Git Ops processes using tools such as Argo CD.
- Build and improve monitoring, logging, tracing, dashboards, and actionable alerting using Prometheus, Grafana, ELK and related observability technologies.
- Improve scalability and workload management using Kubernetes autoscaling technologies such as Karpenter or KEDA.
- Support distributed and event-driven environments, including technologies such as Kafka.
- Develop automation and operational tooling using Python, Bash, Go, or similar languages.
- Strengthen cloud security, resilience, disaster recovery, and production-readiness practices.
- Work directly with enterprise customers when required, including technical troubleshooting, escalations, incident discussions, and explaining infrastructure or reliability issues.
- Collaborate with engineering teams while bringing independent ideas and challenging existing technical approaches where improvements can be made.
- Strong professional experience in Site Reliability Engineering, Cloud Reliability, Platform Engineering, or a comparable production-focused role.
- Senior-level hands-on AWS expertise, with the ability to understand, design, troubleshoot, and improve existing AWS architectures.
- Senior-level Kubernetes experience, ideally operating Amazon EKS in production.
- Strong hands-on experience with Terraform and Infrastructure as Code.
- Proven experience participating in an on-call rotation and personally owning production incidents.
- Demonstrable experience with high-severity incident response, root cause analysis, postmortems, MTTR reduction, and permanent remediation.
- Proven experience communicating directly with external or enterprise customers in a technical capacity, particularly during troubleshooting, escalations, architecture discussions, or production incidents.
- Strong observability experience with Prometheus, Grafana, ELK, or equivalent production observability stacks.
- Strong Linux and cloud networking fundamentals.
- Experience automating operational processes using Python, Bash, Go, or similar scripting/programming languages.
- Experience with CI/CD and Git Ops environments.
- Ability to independently identify infrastructure weaknesses and translate them into practical technical improvements.
- Strong communication skills and professional-level English.
- Comfortable explaining complex technical issues to both engineering teams and customers.
- Hands-on production experience with Argo CD and Git Ops-based deployment workflows.
- Hands-on…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: