Senior AI Platform Engineer
Greater London, London, Greater London, W1B, England, UK
Listed on 2026-08-23
-
Software Development
AI Engineer (Applied/Software), Cloud Engineer - Software, AI Reliability/ Performance Engineer, DevOps
About the job you're considering
Hybrid working: The places that you work from day to day will vary according to your role, your needs, and those of the business; it will be a blend of Company offices, client sites, and your home; noting that you will be unable to work at home 100% of the time.
If you are successfully offered this position, you will go through a series of pre-employment checks, including, identity, nationality (single or dual) or immigration status, employment history going back 3 continuous years, and unspent criminal record check (known as Disclosure and Barring Service)
About usWe build products, not projects: software for insurance claims, payment operations, and health operations, sold to banks, insurers, and health plans. Three product lines run on one shared platform, built by a deliberately small, senior team. Our engineering model is agentic: engineers author the specifications, tooling, evaluation suites, and guardrails, and AI agents do most of the implementation. Humans own every consequential decision, and in our regulated domains some decisions are human-only by design.
Therole
You will build the platform floor that three AI products for banks, insurers, and health plans run on: the model gateway every inference call passes through, the agent runtime that executes planning loops and tool calls, the evaluation infrastructure that gates every release, the guardrail engine, and the shared services (case management, connectors, tenancy, metering) that stop three product teams building the same thing three times.
This is production infrastructure for regulated industries, built by a small senior team with heavy AI leverage. The ambition runs past serving frontier models: we close the loop from production feedback through reinforcement learning and fine-tuning, and train our own LLMs and SLMs where evaluations and economics justify it.
- One or more platform components end to end: model gateway (routing, failover, caching, per-line cost attribution), agent runtime and orchestration, evaluation harness, guardrail engine, tool and agent registry, or shared product services
- The training and adaptation loop: pipelines that turn production traces and evaluation verdicts into reinforcement learning and fine-tuning datasets, and the infrastructure to train, evaluate, and serve New Co-tuned LLMs and SLMs behind the same gates as vendor models
- Reliability, latency, and cost of what you build; platform services carry baselines and you hold them
- The self-service surfaces product engineers use: your components ship with documentation, sane defaults, and no ticket queue
- Co-building with product teams: new shared services start embedded with a product line and graduate to platform services when proven
- Strong production engineering:
Python plus one systems language (Go or Rust welcome), Kubernetes, infrastructure-as-code, and one major cloud - Hands-on experience with LLM infrastructure: a model gateway pattern (such as LiteLLM or in-house), model serving (such as vLLM or managed endpoints), and vector or retrieval systems
- Agent-systems experience: you have built with an agent orchestration framework (such as Lang Graph or first-party agent SDKs) and understand tool calling and the Model Context Protocol (MCP)
- Evaluation engineering: you have built or operated eval harnesses (golden datasets, regression gates in CI, LLM-judge calibration) and can explain why they gate merges
- Post-training fluency: you have fine-tuned or post-trained a model, built the data pipeline behind one, or reproduced techniques from recent papers in production systems
- Daily, hands-on use of AI coding assistants as part of your own development workflow
We care about depth in four or five of these areas more than surface familiarity with all of them.
What sets you apart- Guardrail-engine or AI-observability experience (NeMo Guardrails, Open Telemetry GenAI conventions, Lang Smith, Braintrust, or equivalent)
- Reinforcement learning infrastructure (reward modelling, RLHF or GRPO-class pipelines) or SLM distillation experience
- Multi-tenant SaaS infrastructure: tenancy isolation, metering, usage-based…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).