Cloud Platform Engineer; Senior/Lead
Job in
Farringdon, London, Greater London, W1B, England, UK
Listed on 2026-08-30
Listing for:
Defaqto
Full Time
position Listed on 2026-08-30
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Data Engineering
Job Description & How to Apply Below
Location: Farringdon
Description
Design the Platform Behind Every Product, Data Insight and AI Innovation
As a sole contributor, you'll Lead the design and operation of a modern multi-cloud platform that underpins engineering, data and product teams across the business.
Department: Technology
Location: London
DescriptionYou will:
- Create self-service platforms, golden paths and developer tools that enable teams to ship quickly without managing cloud complexity.
- Build AI-ready infrastructure, guardrails and automation that support both human engineers and autonomous AI agents.
- Champion a zero-trust, secure-by-default approach across identities, networks, workloads and data
- Partner with data teams to optimise cloud data platforms, pipelines and analytics environments that power business-critical insights.
- Design, build and operate our multi-cloud and hybrid-cloud platform across at least two of the top-three providers (AWS, Azure and/or Google Cloud), plus on-prem/hybrid connectivity where needed.
- Build and own an internal developer platform and self-service "golden paths" that make cloud infrastructure feel invisible and commoditised for engineering, data and product teams; and their AI agents.
- Deliver everything as code: infrastructure-as-code, Git Ops, reusable modules, CI/CD pipelines and policy-as-code guardrails.
- Leverage AI agents extensively to automate and optimise platform work - provisioning, cost and performance optimisation, incident response, remediation and documentation.
- Prepare the infrastructure for AI agents as first-class "engineers": safe machine identities, scoped permissions, sandboxes, approval workflows and audit trails so agents can provision and operate infrastructure within tight guardrails.
- Embed a zero-trust security model across identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance.
- Apply SRE practices - SLOs/SLIs, observability, capacity planning, resilience and blameless incident management - to keep the platform reliable and cost-efficient.
- Partner with data engineering to design and optimise data pipelines, data stores and large-scale analytics infrastructure such as Big Query, including query, cost and performance tuning.
- Mentor engineers, set technical direction and champion strong platform and security engineering standards across the organisation.
- Extensive hands-on experience designing, building and operating production cloud infrastructure at senior or lead level.
- Multi-cloud experience across at least two of the top three providers (AWS, Microsoft Azure and Google Cloud), including a recognised professional-level cloud certification for each of those two providers (for example AWS Solutions Architect / Dev Ops Engineer Professional, Azure Solutions Architect / Dev Ops Engineer Expert, or Google Cloud Professional Cloud Architect / Dev Ops Engineer).
- Strong background in modern hybrid-cloud architecture and connecting cloud with on-prem/edge environments.
- Deep infrastructure-as-code and automation skills (e.g. Terraform/Open Tofu, Pulumi, Ansible), Git Ops and CI/CD, plus containers and orchestration (Docker, Kubernetes).
- Proven experience building internal developer platforms, self-service golden paths and platform-as-a-product to abstract away cloud complexity for engineering teams.
- Practical experience using AI agents / LLM-based tooling to automate and optimise infrastructure work, and interest in designing infrastructure that AI agents can operate safely.
- Strong security engineering mindset with hands‑on zero‑trust experience across identity, network, workloads and data — including secrets management, least‑privilege IAM and machine/workload identity.
- Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost‑optimisation practices.
- Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management.
- A third top-tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/CKS).
- Significant data…
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×