Senior AI Reliability Engineer; Platform
Job in
Greater London, London, Greater London, W1B, England, UK
Listed on 2026-07-23
Listing for:
Flatiron Health
Full Time
position Listed on 2026-07-23
Job specializations:
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Location: Greater London
We’re looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. You understand that AI systems fail differently from traditional software – a model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns.
Responsibilities- We’re seeking a Senior AI Reliability Engineer (Platform) to help Flatiron safely and effectively scale AI-enabled workflows across our engineering, product, and business teams
- This role sits at the intersection of data science, platform engineering, AI evaluation, and production reliability
- As Flatiron’s use of AI grows, we need to move beyond experimentation and build the systems, standards, and feedback loops that allow AI workflows to be evaluated, monitored, trusted, and improved over time
- Platform’s strategy is to enable AI adoption without becoming a gatekeeper: building reusable patterns, evaluation infrastructure, observability, and guardrails that help teams move quickly while managing reliability, safety, and cost
- Design, build, and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows
- Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability, drift detection, hallucination risk, retrieval quality, and end‑to‑end workflow behaviour
- Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, alerting, incident response, runbooks, and root‑cause analysis of AI failure modes
- Design orchestration, governance, and guardrails for multi‑agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns
- Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build‑vs‑buy, model selection, and AI adoption decisions
- Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions in a rapidly evolving landscape, collaborating across global teams and participating in on‑call rotations
- Work/life autonomy via flexible work hours and flexible paid time off
- Financial health resources including 1:1 financial advice
- Mental well‑being tools and services
- 401(k) contribution to help you reach your retirement planning goals
- Parental benefits and policies including family‑building care and generous leave
- Path to parenthood programs supporting fertility, adoption and surrogacy
- Travel support for safe healthcare services
- Comfortable operating in ambiguous spaces where the right answer is not always obvious, and motivated by turning emerging AI capabilities into production‑ready systems that teams can actually trust
- Focused on production behaviour and system health, not on pure model research or training
- Senior technical practitioner with experience across data science, machine learning, software engineering, platform engineering, or reliability engineering
- 5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production‑quality systems
- Strong understanding of AI reliability and observability, including logging, tracing, monitoring, drift detection, statistical analysis, uncertainty, alerting, and production system health
- Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG, AI agents, prompt evaluation, and model behaviour
- Knowledge of agentic and multi‑agent systems, including orchestration, state management, tool execution, governance, reliability, human‑in‑the‑loop controls, and selecting the appropriate level of AI autonomy for a given problem
- Experience with modern cloud and ML infrastructure, including AWS, containers, Kubernetes,…
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×