×
Register Here to Apply for Jobs or Post Jobs. X

Principal Software Reliability Engineer

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Entrust Datacard
Full Time position
Listed on 2026-02-13
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing, SRE/Site Reliability, Cybersecurity
Salary/Wage Range or Industry Benchmark: 80000 - 100000 GBP Yearly GBP 80000.00 100000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Join us at Entrust

At Entrust, we’re shaping the future of identity centric security solutions. From our comprehensive portfolio of solutions to our flexible, global workplace, we empower careers, foster collaboration, and build solutions that help keep the world moving safely .

Get to Know Us

Headquartered in Minnesota, Entrust is an industry leader in identity-centric security solutions, serving over 150 countries with cutting-edge, scalable technologies. But our secret weapon? Our people. It’s the curiosity, dedication, and innovation that drive our success and help us anticipate the future.

About This Role

This is a Product Reliability position, not an infrastructure SRE role. Our Dev Ops team manages the infrastructure platform; this role focuses on application and service-level reliability, working directly with product engineers.

This is the first role of its kind in product engineering. Reporting to the VP of Product Engineering for Consumer Identity, you’ll drive reliability efforts across the team: defining the roadmap, prioritizing initiatives, and partnering with engineering directors and senior ICs to deliver them.

Why Join Us
  • Greenfield opportunity:
    You’ll define Product Reliability as a discipline here. Build the playbook, not inherit one.
  • High-impact domain:
    Consumer Identity powers identity verification and biometric authentication for some of the world’s largest financial institutions. Our reliability directly impacts fraud prevention and customer onboarding at scale.
  • Real authority:
    Direct line to VP Engineering, budget for tooling, seat at architecture council and service reviews.
  • Strong foundation:
    We’re not firefighting. 99.98% uptime means you’re optimizing, not triaging chaos.
  • Technical depth:
    Work across ML pipelines, computer vision systems, and mobile SDKs (not just YAML and dashboards).
  • Ownership culture:
    Engineers own their services end-to-end; you’ll amplify that, not replace it.
Experience Level Staff SRE
  • 8+ years in software engineering
  • 4+ years in reliability/SRE
  • Drives reliability initiatives across multiple teams; hands-on with complex systems
Principal SRE
  • 15+ years in software engineering
  • 6+ years in reliability/SRE
  • Sets technical direction org-wide; influences business-unit-level reliability strategy

We’re open to either level. Scope and compensation will match your experience. Principal candidates should demonstrate cross-org impact and a track record of building reliability programs from scratch.

Current State Incident Analysis (2020–2025)
  • Postmortem volume peaked in 2023, down 48% since then despite increased release cadence
  • P0-to-P1+ ratio remains stable despite lower overall incident volume
  • 65% change-induced incidents (deployments, migrations, config changes); 35% organic (third-party outages, expirations, attacks)
  • Change-induced ratio improved modestly: 69% → 62%
  • Detection time: 35 min → 18 min
  • Customer-first detection: 40% → 22%
Availability Targets
  • 2024 & 2025 average uptime: 99.98% (as available in our public status page)
  • Goal:
    Consistent 99.99% (four nines) average uptime, SLO breach reductions
System Simplification
  • We’re reducing system complexity to narrow the reliability target area:
  • Microservices (K8s deployments/rollouts) reduced 29% from peak, with further cuts planned for 2026
  • Goal:
    Smaller footprint, higher reliability, lower cost for new regions
Role Objectives
  • Primary goal:
    Improve release safety, reduce releases that cause downtime or SLO degradation.
  • We already have foundational systems in place:
  • Automated test coverage and crowd testing
  • A/B testing and dark canaries
  • Progressive rollouts (infrastructure and application level)
  • Back-testing against historical data
  • To consistently exceed four nines, we need to mature these systems and build new capabilities.
Ideal Candidate Profile Mindset
  • Passionate about reliability as a discipline, not just a checkbox
  • Focused on reliability, not product features, but willing to learn the product to understand impact
  • Hands-on: eager to build tooling and systems
  • Pragmatic about balancing reliability with development velocity
Required Skills Software Engineering
  • Strong software engineering in at least one of our backend languages (Python, Ruby,…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)

Job Posting Language
Employment Category
Education (minimum level)
Filters
Education Level
Experience Level (years)
Posted in last:
Salary