×
Register Here to Apply for Jobs or Post Jobs. X

Senior Cloud Engineer, AI Platform SRE

Job in City of Edinburgh, Edinburgh, City of Edinburgh Area, EH1, Scotland, UK
Listing for: CreateFuture
Full Time position
Listed on 2026-08-13
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 90000 - 130000 GBP Yearly GBP 90000.00 130000.00 YEAR
Job Description & How to Apply Below
Location: City of Edinburgh

Working at Create Future
Create Future is an AI-native consulting partner where people do work that matters and are supported to do it well. We work alongside organisations such as Pay Pal, adidas, Nat West, Fan Duel and Money Saving Expert, building digital products and services that make a difference while always putting people first.

We’re a team of creators. We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people. We work side by side with our clients, challenging what’s not working and helping them to build the future. Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.

  • 35 days leave (including bank holidays).
  • Private medical insurance.
  • Enhanced parental and adoption leave.
  • 40 hours of paid learning and development.

Join us on our journey. Let’s create tomorrow, together, today.

About the role and team

AI platforms are still mostly run like prototypes. There's usually a Kubernetes cluster somebody set up in a hurry, no SLOs, no runbooks, and a cost line nobody can explain. We're engaging a Senior Cloud Engineer to bring proper SRE discipline to a major AI platform programme with a client in a high-traffic, heavily regulated consumer sector. You'll own how the platform runs.

That means the Kubernetes infrastructure behind model serving, agent orchestration and batch inference. It means the CI/CD pipelines that ship models, agents, tools and prompt changes. It means the SLOs and on-call practice that make reliability an asset commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops.

You'll be the operational conscience of the programme. When a new capability is about to ship without limits, monitoring or a runbook, you're the person who says so. This is a role for someone who finds that work satisfying rather than thankless, and who wants to do it on a platform where the SRE patterns are still being written.

What you'll be doing Technical Delivery & Implementation
  • Kubernetes for AI workloads:
    Design, build and operate the Kubernetes infrastructure behind model serving, agent orchestration and batch inference, defined in Terraform and deployed through Git Ops.
  • CI/CD for non-deterministic systems:
    Build deployment pipelines suited to AI workloads, covering model rollouts, agent and tool updates, and prompt and configuration changes, with automated eval and regression checks before promotion.
  • Reliability engineering:
    Define and track SLOs and SLAs for platform services, run incident response and root cause analysis, and write post-mortems people will read.
  • On-call and runbooks:
    Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am.
  • AI-specific observability:
    Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users do.
  • Fin Ops:
    Embed cost estimation, anomaly detection and resource optimisation into the pipeline as a default rather than a monthly review.
  • Security posture:
    Apply zero-trust networking, least-privilege access and secrets discipline to a platform handling sensitive data across multiple third-party model providers.
Client Delivery & Stakeholder Management
  • Operability from day one:
    Work with the inference control plane, evaluation and knowledge platform work streams so new capabilities arrive with monitoring, limits and a runbook, instead of acquiring them after the first incident.
  • Technical advisory:
    Guide the client's architecture and AWS service choices on this platform, and build the working relationships that let that guidance land.
  • Cost accountability:
    Understand the programme's budget context, build the cost‑benefit case for infrastructure decisions, and be able to justify spend to a finance audience.
  • Risk and delivery:
    Spot emerging operational bottlenecks in the client's estate, raise them early, and plan and estimate your own stream accurately.
  • Stakeholder communication:
    Explain…
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary