AI Platform Engineer
Listed on 2026-09-01
-
Software Development
DevOps, Cloud Engineer - Software, AI Engineer (Applied/Software)
Salary: up to £115k (+ very generous early-stage equity, up to ~£90k)
Location: Central London, EC1 (3 office day/week)
This London startup is building a new intelligence layer designed to bring more context and security to digital payments. Their technology analyses transactions in real time, gathering signals from multiple sources to determine whether a payment is legitimate or potentially fraudulent.
The platform combines distributed data systems, real-time investigations and AI-driven decisioning to help financial institutions detect scams while allowing legitimate payments to flow without unnecessary friction. Within 2 years of being founded, they're working with most Tier 1 banks and payment providers in the UK - and are about to launch in the US!
Hiring an AI Platform Engineer to help build the foundations that let a fast-moving engineering team ship safely to some of the country's largest financial institutions. Production is increasingly powered by non-deterministic AI agents, so this isn't a "keep the lights on" role - you'll be defining what good looks like for platform and infrastructure culture, not inheriting someone else's playbook.
Key responsibilities:- Owning the reliability and operability of production systems - monitoring, alerting, incident response and post-incident learning
- Maintaining and evolving the infrastructure-as-code estate, making it easy to ship safely and hard to ship dangerously
- Securing infrastructure defaults so the easy path is the safe path
- Designing observability across the stack - metrics, traces, logs, dashboards and alerts - including for AI agent behaviour, where "correct" isn't always the same twice
- Driving incident response maturity from detection through resolution to follow-up
- Building platform capabilities that unblock engineering teams - deployment pipelines, release tooling, developer experience
- Building the guardrails and automation (including AI-assisted triage and response) that let the wider team move fast without breaking things
- Supporting the platform's evolution from stability through to scalability as the company grows
- At least 3-4 years in platform/Dev Ops/SRE, ideally with genuine ownership at an early-stage startup - comfortable with ambiguity
- Strong Terraform/IaC experience on real production infrastructure
- Deep cloud infrastructure experience (GCP a strong plus)
- Proven observability track record - monitoring, alerting and dashboards for distributed systems
- Software Engineering background / proficiency in Python or Go at a decent level
- Incident response experience - on-call, running incidents, and building the processes that make both better
- Security foundations - least-privilege access, secrets management, secure-by-default infrastructure
- CI/CD experience - deployment pipelines teams trust, with a focus on deploy velocity and rollback safety
- Genuine hands-on exposure to agentic AI frameworks (Lang Chain, Lang Graph or similar) within the last ~12 months - not just conceptual awareness
- Fintech or other regulated-industry experience
- Experience building observability/reliability for non-deterministic or ML-powered systems specifically
- Exposure to compliance frameworks (ISO 27001, SOC
2) - Experience with workflow orchestration engines (Temporal or similar)
- A track record of bootstrapping a platform/Dev Ops function from scratch
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: