Reliability Engineer 4 (Observability Specialist
Listed on 2026-08-21
-
IT/Tech
SRE/Site Reliability, IT Support
Senior-Level Reliability Engineer Specializing in Observability
This role partners closely with product owners, application engineering teams, SRE teams, and business stakeholders to translate customer journeys and business outcomes into measurable reliability objectives. The role establishes and maintains best-practice processes for documenting, governing, reviewing, and improving user journeys, SLIs, SLOs, synthetic monitoring, dashboards, alerts, telemetry standards, and related observability assets. The engineer provides senior technical guidance, identifies observability gaps through incident and performance analysis, drives continuous improvement, and helps ensure teams have the data, processes, and operating discipline needed to detect issues earlier, reduce customer impact, and improve overall service reliability.
Responsibilities include:
- Leading the definition, documentation, implementation, and continuous improvement of Observability across Critical Customer Journeys, ensuring alignment between Observability Strategy, business outcomes, and reliability objectives.
- Designing, implementing, and governing Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability metrics for enterprise applications and services.
- Establishing and maintaining Observability Governance Frameworks for Dashboards, Alerts, Synthetic Monitoring, Telemetry Standards, and lifecycle management of observability assets.
- Translating business and technical requirements into scalable Observability Architectures, including Instrumentation Standards, Monitoring Strategies, Tagging Frameworks, and Alerting Models.
- Partnering with Product Owners, Application Engineering, Site Reliability Engineering (SRE), and Operations Teams to ensure applications are production-ready and fully instrumented for reliability measurement.
- Developing and maintaining executive and operational Service Health Dashboards that provide insights into Availability, Latency, Customer Impact, Dependency Performance, and SLO Compliance.
- Analyzing Telemetry Data, Incident Trends, Problem Records, and Alert Performance to identify observability gaps, reduce alert fatigue, and improve detection accuracy.
- Providing technical leadership and mentorship on Distributed Tracing, Logging, Metrics Collection, Synthetic Testing, Application Performance Monitoring, Real User Monitoring, Monitoring Design Patterns, and Alert Governance Best Practices.
Basic Qualifications:
- Bachelor's degree, or equivalent work experience.
- Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.
Preferred Skills/
Experience:
- Expertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering.
- Strong knowledge of SLIs, SLOs, Error Budgets, and Customer Journey Monitoring.
- Demonstrated ability to understand stakeholder needs and guide the development of reliability requirements for large, complex multi-system products.
- Hands-on experience with APM, RUM, synthetics, monitoring, logging, tracing, and telemetry frameworks.
- Proficiency with Datadog, Dynatrace, Splunk, Grafana, Prometheus, New Relic, Elastic, or Open Telemetry.
- Experience building, standardizing, and tuning operational dashboards and actionable alerts that communicate service health, customer impact, dependency health, performance trends, failure conditions, severity, ownership, routing, and runbook linkage.
- Strong understanding of distributed systems, microservices, cloud platforms, and Kubernetes.
- Ability to leverage incident analysis, RCA, and performance data to drive reliability improvements.
- Excellent stakeholder management, communication, and technical leadership skills.
Location expectations:
This role requires working from a U.S. Bank location three (3) or more days per week.
Benefits:
- Healthcare (medical, dental, vision)
- Basic term and optional term life insurance
- Short-term and long-term disability
- Pregnancy disability and parental leave
- 401(k) and employer-funded retirement plan
- Paid vacation (from two to five weeks depending on salary grade and tenure)
- Up to 11 paid holiday opportunities
- Adoption assistance
- Sick and Safe Leave accruals of one hour for every 30 worked, up to 80 hours per calendar year unless otherwise provided by law
U.S. Bank is an equal opportunity employer. We consider all qualified applicants without regard to race, religion, color, sex, national origin, age, sexual orientation, gender identity, disability or veteran status, and other factors protected under applicable law.
Pay Range: $ - $
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).