×
Register Here to Apply for Jobs or Post Jobs. X

Principal Platform Engineer

Job in London, Greater London, W1B, England, UK
Listing for: black.ai
Per diem position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 70000 - 110000 GBP Yearly GBP 70000.00 110000.00 YEAR
Job Description & How to Apply Below

Who is Heidi?

Heidi is building an AI Care Partner that supports clinicians every step of the way, from documentation to delivery of care.

We exist to double healthcare’s capacity while keeping care deeply human. In 18 months, Heidi has returned more than 18 million hours to clinicians and supported over 73 million patient visits. Today, more than two million patient visits each week are powered by Heidi across 116 countries and over 110 languages.

Founded by clinicians, Heidi brings together clinicians, engineers, designers, scientists, creatives, and mathematicians, working with a shared purpose: to strengthen the human connection at the heart of healthcare.

Backed by nearly $100 million in total funding, Heidi is expanding across the USA, UK, Canada, and Europe, partnering with major health systems including the NHS, Beth Israel Lahey Health, Maine General, and Monash Health, among others.

We move quickly where it matters and stay grounded in what’s proven, shaping healthcare’s next era. Ready for the challenge?

The Role

This role sits in the core Platform/SRE team that owns production. You’ll work directly on incident response, on‑call, system reliability, and day‑to‑day operations for Heidi’s platform.

We’re open to candidates who are strong mid‑level SREs ready to take on more ownership, as well as senior SREs who enjoy being hands‑on in operations. The role is intentionally ops‑heavy and focused on keeping real systems healthy in production.

What you’ll do
  • Participate in on‑call and incident response: Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end‑to‑end.

  • Improve operational reliability: Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.

  • Own parts of the production environment: Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.

  • Strengthen observability: Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.

  • Reduce operational toil: Automate repetitive tasks, simplify runbooks, and improve tooling to make on‑call and day‑to‑day operations easier and safer.

  • Support safe change: Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.

  • Contribute to operational practices: Write and maintain runbooks, participate in blameless post‑mortems, and help improve incident response processes over time.

  • Collaborate closely with engineers: Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.

What we’re looking for
  • 3–6+ years in SRE, Dev Ops, Platform, or operations‑heavy engineering roles.

  • Experience supporting production systems and participating in on‑call rotations.

  • Comfortable debugging live systems under pressure.

  • Experience operating cloud infrastructure (AWS preferred).

  • Working knowledge of Kubernetes and containerised workloads.

  • Infrastructure as Code experience (Terraform or similar).

  • Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc).

  • Scripting or automation experience (Python, Bash, or similar).

Nice to have:

  • Experience leading incidents or mentoring others during on‑call.

  • Experience in regulated or security‑sensitive environments.

  • Familiarity with databases, queues, and caches in production.

  • Interest in reliability practices such as SLOs, error budgets, and capacity planning.

How We Work
  • We own production: The Platform/SRE team is responsible for reliability and incident response.

  • Incidents are blameless: We focus on learning and improving systems, not assigning fault.

  • Practical over perfect: We prioritise improvements that reduce real operational pain.

  • Calm under pressure: Clear thinking and communication matter during incidents.

What do we believe in?

Heidi builds for the future of healthcare, not just the next quarter, and our goals are ambitious because the world’s health…

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary