×
Register Here to Apply for Jobs or Post Jobs. X

SRE Observability Technical Lead - Vice President

Job in Dundonald, County Down, BT16, Northern Ireland, UK
Listing for: Citi
Full Time position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 50000 - 70000 GBP Yearly GBP 50000.00 70000.00 YEAR
Job Description & How to Apply Below

Engineer the future of global finance. At Citi, our Tech team doesn’t just support finance – we are helping to redefine it. Every day, $5 trillion crosses through our network. We do business in 180+ countries operating at a scale few can match. From deploying advanced AI to helping shape global markets, we build systems that matter.

Help to join a team where your work helps influence economies, your ideas can drive innovation and outcomes, and your growth is backed by mentorship, continuous learning and flexibility with potential hybrid work opportunities. Help solve real-world challenges that touch millions and get the opportunity to build the future of finance with Citi Tech.

The SRE Observability Specialist is a hands‑on expert, delivering the future of Observability across Services Technology. This role is a part of a central SRE enablement team within Services Production, working closely with SREs, developers, and platform teams to embed telemetry, implement SLOs, and build meaningful visualizations for key production flows — particularly in critical Payments Business.

Key Responsibilities
  • Define the roadmap for Engineering enablers for Project Orion team aligned with enterprise reliability and SRE Services organization goals.
  • Translate Organization strategy into an actionable delivery plan in partnership with Services Products, Operations & Engineering function, delivering incremental, high‑value milestones.
  • Understand Critical Business Services functional scope and translate into End‑to‑End monitoring solutions.
  • Deliver against the observability roadmap for Services Technology by building scalable, reusable telemetry solutions.
  • Periodic review and analyze application monitoring TOIL and collaborate with stakeholders and remediate them as per organization goal.
  • Create and maintain dashboards and visualizations for critical client journeys, including real‑time flows across Payments.
  • Guide line‑of‑business teams in implementing SLIs/SLOs, golden signals, and effective alerting to support operational excellence.
  • Support integration and adoption of observability tooling across on‑prem, public cloud (AWS/GCP), and containerized environments (ECS, Kubernetes).
  • Customize shared dashboards and observability components in partnership with CTI and other central Engineering functions, ensuring usability and flexibility.
  • Provide technical support and implementation guidance to SREs and developers facing integration or tooling challenges.
  • Effectively manage the observability book of work for Services Technology and drive initiatives to reduce MTTD and improve recovery outcomes.
  • Serve as a key connection point between line‑of‑business SREs and central infrastructure functions by gathering tooling feedback, surfacing systemic issues, and influencing platform enhancements via the Services Observability Forum.
  • Stay current with observability trends, including AI/ML‑driven insights, anomaly detection, and emerging OSS practices, and assess their applicability.
  • Maintain strong knowledge of observability platform features and vendor offerings to advise teams and maximize the value of tooling investments.
  • Foster AI adoption by building use cases performed by Orion L1 Functions and remediation using Citi AI tech stack.
Qualifications
  • Experience in SRE, Observability Engineering, or platform infrastructure roles focused on operational telemetry.
  • Hands‑on experience in observability tools and stacks such as Grafana, Prometheus, Open Telemetry, ELK, Splunk, and similar platforms.
  • Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high‑availability environments.
  • Proven ability to troubleshoot integration issues and support observability across hybrid platforms (on‑prem, cloud, containers).
  • Experience building dashboards aligned to business outcomes and incident workflows, especially in critical flows such as payments.
  • Familiarity with modern observability tooling ecosystems, including AI/ML capabilities, trace correlation, baselining, and alert tuning.
  • Strong interpersonal and collaboration skills – able to operate across federated engineering teams and central infrastructure groups.
  • Experience in…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary