SRE Observability Technical Lead - Vice President
Listed on 2026-07-21
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Engineer the future of global finance. At Citi, our Tech team doesn’t just support finance – we are helping to redefine it. Every day, $5 trillion crosses through our network. We do business in 180+ countries operating at a scale few can match. From deploying advanced AI to helping shape global markets, we build systems that matter. Look to join a team where your work helps influence economies, your ideas can drive innovation and outcomes, and your growth is backed by mentorship, continuous learning and flexibility with potential hybrid work opportunities.
Help solve real-world challenges that touch millions and get the opportunity to build the future of finance with Citi Tech.
The SRE Observability Specialist is a hands‑on expert, delivering the future of Observability across Services Technology. This role is a part of a central SRE enablement team within Services Production, working closely with SREs, developers, and platform teams to embed telemetry, implement SLOs, and build meaningful visualizations for key production flows – particularly in critical Payments Business.
This role requires providing technological support solution for Function called Project Orion which provides End‑to‑End payment monitoring like Building an End‑to‑End payments Dashboard, Toil Reduction, Transformation of legacy monitoring into observability based monitoring solution, requires good understanding of different Payments Taxonomy (ACH, Wires, Instant Payments, etc.). Strong commercial awareness, technical credibility, and excellent communication skills are essential to negotiate internally, influence peers, and drive change.
Some external communication may be necessary.
- Define the roadmap for Engineering enablers for Project Orion team aligned with enterprise reliability and SRE Services organization goals.
- Translate Organization strategy into an actionable delivery plan in partnership with Services Products, Operations & Engineering function, delivering incremental, high‑value milestones.
- Understand Critical Business Services functional scope and translate into End‑to‑End monitoring solutions.
- Deliver against the observability roadmap for Services Technology by building scalable, reusable telemetry solutions.
- Periodic review and analyze application monitoring TOIL and collaborate with stakeholders and remediate them as per organization goal.
- Create and maintain dashboards and visualizations for critical client journeys, including real‑time flows across Payments.
- Guide line‑of‑business teams in implementing SLIs/SLOs, golden signals, and effective alerting to support operational excellence.
- Support integration and adoption of observability tooling across on‑prem, public cloud (AWS/GCP), and containerized environments (ECS, Kubernetes).
- Customize shared dashboards and observability components in partnership with CTI and other central Engineering functions, ensuring usability and flexibility.
- Provide technical support and implementation guidance to SREs and developers facing integration or tooling challenges.
- Effectively manage the observability book of work for Services Technology and drive initiatives to reduce MTTD and improve recovery outcomes.
- Serve as a key connection point between line‑of‑business SREs and central infrastructure functions by gathering tooling feedback, surfacing systemic issues, and influencing platform enhancements via the Services Observability Forum.
- Stay current with observability trends, including AI/ML‑driven insights, anomaly detection, and emerging OSS practices, and assess their applicability.
- Maintain strong knowledge of observability platform features and vendor offerings to advise teams and maximize the value of tooling investments.
- Foster AI adoption by building use cases performed by Orion L1 Functions and remediation using Citi AI tech stack.
- Experience in SRE, Observability Engineering, or platform infrastructure roles focused on operational telemetry.
- Bachelor's degree (computer science or related fields) or equivalent experience in building scalable solutions to improve the service reliability and/or increase productivity and efficiency.
- Hands‑on…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: