More jobs:
Lead DevOps Engineer
Job in
London, Greater London, W1B, England, UK
Listed on 2026-08-12
Listing for:
Collinson Group
Full Time
position Listed on 2026-08-12
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Cybersecurity, IT Project Manager
Job Description & How to Apply Below
Travellers can access a network of 1,500+ lounges and travel experiences, including dining, retail, sleep and spa, in over 650 airports in 148 countries, helping to elevate the journey into something special. We work with the world’s leading payment networks, over 1,400 banks, 90 airlines and 20 hotel groups worldwide.
We have been bringing innovation to the market since inception – from launching the first independent global VIP lounge access Programme, Priority Pass to being the first to sell direct travel insurance in the UK through Columbus Direct and creating the first loyalty agency of its kind in the travel sector with ICLP. Today we still invest heavily in innovation to ensure that we continue to deliver superior customer experiences.
Key clients include Mastercard, American Express, Cathay Pacific, British Airways, LATAM, Flying Blue, Accor, Easy Jet, HSBC, Chase, HDFC.Our mission is focused on doing good beyond profit, which for us means we seek out opportunities for our people to share in our success and that we give back to the communities and people within which we work.
Never short of ambition, the success of our business is delivered through the diverse and talented team of over 2,200 global colleagues.
Lead Dev Ops Engineer
The Role As Lead Dev Ops Engineer, you'll own the design, reliability, scalability, security, and cost efficiency of our production platform - a zero-downtime, multi-region AWS environment running Kubernetes and serverless workloads and centralised AI enablement layers 'll set the technical direction for the Dev Ops team, driving and documenting the platform standards, practices, and tooling that champion Dev Ex, eases friction and keep our systems resilient and secure.
This is a hands-on leadership role. You’ll architect and solve complex infrastructure challenges while mentoring and growing the engineers around you. You'll own our platform SLAs, serve as the platform advocate in cross-functional decisions, and act as the core authority on resilience, cost, security, and enablement.
Key Responsibilities Platform Ownership
- Own the end-to-end reliability, stability, and performance of our production platform. Enforce SLAs across availability, security posture, and resilience. Drive the measures, thresholds, and incident response standards the team operates to.
Zero-Downtime Operations
- Run and maintain a multi-region, zero-downtime platform. Own deployment strategies (blue/green, canary), traffic management, and failover patterns that ensure continuous availability under change and failure.
Security & Compliance
- Own the platform's security posture end to end: networking controls, IAM, KMS, secrets management, vulnerability management, and compliance against frameworks such as PCI DSS v4 and CIS Benchmarks. Work closely with security teams and act as the Dev Ops authority on audit and risk.
DR & Business Continuity
- Participate in disaster recovery strategy, automation, execution and continuous improvement.
TCO & Cost Management
- Drive the strategy on platform total cost of ownership, providing teams with clear cost-allocation visibility. Drive rightsizing, tagging strategy and architectural decisions that balance cost against reliability and developer experience.
Infrastructure as Code & CI/CD
- Set and maintain the standards for IaC (Terraform, Terragrunt). Ensure deployments are secure, fast, and auditable. Raise the bar on automation so routine operational tasks are eliminated, not managed.
Observability
- Own the observability strategy across the platform using Datadog. Define what good looks like: SLO/SLA dashboards, alerting thresholds, runbooks, and the feedback loops that let teams act on signals before users feel them.
Team Leadership & Mentoring
- Lead and grow a team of Dev Ops engineers. Set technical direction, conduct design reviews, and create an environment where engineers take ownership. Actively mentor and develop less experienced engineers, helping them move from task execution to platform thinking.
AI & Automation
- Lead the team's adoption of AI-assisted operations - embedding AI into CI/CD pipelines, runbook automation, incident response, and developer tooling. Define where AI adds real leverage and own its safe integration into the platform.
Your Experience Deep expertise with AWS at scale: EKS (Automode), Lambda, EC2, RDS/Aurora, multi-region networking (VPC + Endpoints, Transit Gateway, Network Firewall, WAF, API Gateway), security services (IAM, KMS, Secrets Manager, SCPs, Security Hub, Guard Duty) and AI tooling (Bedrock…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×