×
Register Here to Apply for Jobs or Post Jobs. X

Observability Tech Lead

Job in London, Greater London, W1B, England, UK
Listing for: Trade Republic
Full Time position
Listed on 2026-09-12
Job specializations:
  • Software Development
    DevOps, Backend Developer, Cloud Engineer - Software, Software Architect
Job Description & How to Apply Below

Please note that these positions are based in London Berlin or Paris relocation support is provided if required.



THE BEST WORK OF YOUR CAREER

Trade Republic is the largest savings platform in Europe - we operate in 18 countries serving 10 million customers who trusted us with over 150B in assets. But were striving for more.

We have a bold mission to empower everyone to build wealth with easy safe and free access to financial systems. You will have the opportunity to grow your career by collaborating with a team of outstanding talents and state of the art technology to build a lasting positive future for millions. ing talents and state of the art technology to build a lasting positive future for millions.

ABOUT PLATFORM ENGINEERING

Platform Engineering is the backbone of Trade Republics engineering velocity. Our mission is to build scalable platforms for a Europe-scale bank serving internal engineers and building in-house control planes for managing the banks infrastructure. Were a 50-person Platform team focused on one thing: enabling product engineers to move fast and operate autonomously by default.

We build self-service platforms golden paths and opinionated tooling so that over 400 engineers can ship with confidence. From Kubernetes fleet management and CI/CD to an internal Developer Hub built on Backstage our work underpins every trade savings plan and card payment that flows through the platform.

THE OBSERVABILITY JOURNEY

In 2023 we made a decisive move: we replaced our observability-as-a-service provider with a fully self-hosted observability stack giving us complete control over cost data residency and the developer experience around telemetry. Today our stack spans the full LGTM suite Grafana Mimir Loki and Tempo alongside Victoria Metrics self-hosted Sentry Grafana Alloy as our telemetry collector and Open Telemetry as the instrumentation standard.

We use Pyrra for SLO tracking and are building toward a unified service health dashboard powered by error budgets and burn-rate alerting.

Telemetry is the backbone of how we operate a bank at scale ingesting over 100 million samples serving 400 services and capturing end-to-end traces from clients through services to system dependencies. Every trade every card payment depends on our ability to see measure and respond to whats happening in production. Weve proven the architecture works. Now were building a dedicated in-house observability team to take it to the next level: stabilise and harden the platform drive down cost-per-signal and build the golden path for observability where 100% of components ship with production-grade telemetry because the best thing to do is the easiest thing to do.

WHAT YOULL BE DOING

  • Build and evolve the observability platform: Design and operate large-scale telemetry pipelines while continuously improving core components with a strong focus on automation reliability and developer experience.
  • Build for scale design for cost: Architect high-throughput telemetry systems with sampling strategies data tiering and retention policies that balance signal fidelity with infrastructure cost at scale.
  • Make production observable by default: Define and implement observability and reliability standards SLOs error budgets and low-noise alerting and actively support engineering teams in adopting them making doing the right thing effortless.
  • Own the platform end to end: Participate in the on-call rotation for the observability platform ensuring full end-to-end ownership of the systems you build and operate.
  • Own the direction and drive it forward: Define long-term observability direction drive cross-team initiatives from kickoff to delivery and align observability strategy with broader engineering reliability and business goals.

WHAT WERE LOOKING FOR

  • 5 years of experience in observability platform engineering or a related SRE/infrastructure discipline.
  • We are hiring from senior to staff level so whether you have a strong foundation and are ready for more ownership or you have been leading observability strategy for large-scale systems for many years we would love to hear from you.
  • Deep hands-on expertise with the observability stack Prometheus Open Telemetry Grafana or equivalent ds-on experience with Mimir Loki and Tempo architectures is a strong benefit.
  • Proven ability to design and operate high-throughput telemetry pipelines across distributed multi-cloud environments.
  • Strong command of SLO-based reliability practices error budgets burn-rate alerting and incident…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary