Senior Platform Engineer
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability
Location: Greater London
This role is with one of Dex's trusted partner companies. We work closely with their teams to truly understand their culture, goals, and what they're looking for, so we can match you with the right opportunity and give you context about the role before you commit to a process.
The roleBanks need to screen payments in real time, intercepting scams while letting genuine transactions flow. This company builds the intelligence layer that makes that possible, increasingly powered by AI agents whose behavior isn't static. They're live with UK banks, showing contextual warnings significantly more effective than generic screens, and are backed by leading investors. This isn't a 'keep the lights on' role.
You'll own the reliability and operability of production systems, defining what good looks like for Platform at a company shipping to large banks with a small, high-calibre engineering team. Expect to balance enterprise resilience with startup velocity, building observability for non-deterministic AI behavior and driving incident response maturity from first detection to post-incident learning.
- Own the reliability and operability of production systems, including monitoring, alerting, and incident response.
- Build observability specifically for AI-agent behavior, where traditional methods fall short.
- Maintain and evolve infrastructure-as-code (Terraform on GCP), making the safe and secure path the easy default.
- Drive incident response maturity, from detection through to follow-up and post-incident learning.
- Implement deployment pipelines, release tooling, and developer experience improvements to unblock the engineering team.
- Strong, hands-on Terraform experience running real production infrastructure; GCP experience is a plus.
- Built monitoring, alerting, and dashboards for distributed systems, with exposure to non-deterministic or ML-powered systems.
- Experience on-call, running incidents calmly, and writing post-mortems that prevent recurrence.
- Understanding of security fundamentals: least-privilege access, secrets management, and secure-by-default infrastructure design.
- 5+ years working in early-stage companies, comfortable defining best practices where no playbook exists.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: