×
Register Here to Apply for Jobs or Post Jobs. X

Graduate Engineer: AI Tooling and Site Reliability

Job in Cardiff, Cardiff City Area, CF10, Wales, UK
Listing for: Critical Cloud Limited
Full Time position
Listed on 2026-06-20
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), AWS
Salary/Wage Range or Industry Benchmark: 40000 GBP Yearly GBP 40000.00 YEAR
Job Description & How to Apply Below

We're building an internal AI platform from scratch, the tooling that will define how Critical Cloud operates as we scale across Europe. This isn't a rotation or a shadow programme. From week one you'll be shipping real tooling and operating real production environments for real customers. The two tracks exist because they make each other better. That's the design.

About the Role

This isn't a rotation programme. From week one, you'll contribute to both tracks: shipping AI tooling that helps us run cloud operations better, and operating real production infrastructure for real customers. Two disciplines, one engineer, no siloes.

Critical Cloud is the world's first "Powered by Datadog" accredited MSP, a Datadog-native cloud MSP built for European tech‑led SMBs. We're building an internal AI platform (the Critical Cloud Platform) to automate and augment how we operate customer environments. This role sits at the centre of that programme.

Half your time will be engineering AI‑assisted tooling: LLM integrations, agents, and automation workflows that reduce toil and improve our operational quality. The other half will be hands‑on SRE work: monitoring, incident support, infrastructure‑as‑code, and customer‑facing operations. Each half makes you better at the other.

What You’ll Do AI Tooling Track
  • Build and iterate on AI‑assisted automation workflows using LLM APIs (Claude, OpenAI) integrated with cloud and observability tooling
  • Develop tooling for automated infrastructure discovery, customer onboarding, and operational runbook generation
  • Contribute to the Critical Cloud Platform: our internal AI governance framework and agent operating model
  • Design and implement MCP (Model Context Protocol) integrations connecting AI agents to Datadog, AWS, and Azure APIs
  • Write evaluation harnesses and regression tests to keep AI tool output reliable and auditable
  • Document AI system behaviour against our constitutional operating framework and ISO 27001 controls
Site Reliability Track
  • Monitor and triage alerts across customer AWS and Azure environments using Datadog as the primary observability platform
  • Support incident response workflows and contribute to postmortem documentation alongside the SRE team
  • Support Datadog onboarding for new customers: instrumentation, dashboards, monitors, and SLO configuration
  • Write and maintain Terraform modules for infrastructure provisioning and change management
  • Produce and maintain operational runbooks, escalation guides, and change records to ISO 27001 standards
  • Contribute SRE context back into AI tooling: you'll know what's worth automating because you've done it manually
Requirements
  • A degree in Computer Science, Software Engineering, or a related technical field (2:1 or above)
  • Solid Python: comfortable writing scripts, working with APIs, and handling structured data
  • Familiarity with cloud fundamentals (AWS or Azure), ideally through coursework, personal projects, or placement
  • Experience consuming REST APIs or LLM APIs, whether through a project, dissertation, or side work
  • Clear written communication: you'll be writing docs and talking to customers
Nice to Have
  • Hands‑on LLM work: prompt engineering, tool use, agent frameworks, or evaluation pipelines
  • Terraform or any IaC tooling (even tutorials count)
  • Datadog experience, even a free tier account you've played with
  • Kubernetes or containerised workload exposure
  • Any cloud or AI certification (AWS, Azure, Google, or Datadog)
  • A Git Hub profile with something worth showing us
AI & Automation

Claude / Anthropic API – Primary LLM platform

Datadog – Core observability platform

AWS – Primary cloud, multi‑account

Azure – Secondary cloud workloads

Terraform – Infrastructure as code

Git Hub Actions – CI/CD pipelines

Start

Year 1–2

Year 2–3

Engineer II – Specialise or Broaden

Year 3+

Senior / Lead – Platform or SRE

Benefits
  • 25 days holiday + bank holidays plus a paid day off in your birthday month, taken in the month it falls
  • Holiday grows with tenure: +1 day per year after your second work anniversary, up to 28 days total
  • Enhanced maternity pay: 26 weeks at your full basic salary
  • Enhanced paternity pay: 2 weeks at your full basic salary
  • Datadog, AWS, Azure, and AI tooling certifications…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary