×
Register Here to Apply for Jobs or Post Jobs. X

Reliability Engineer ; Observability Specialist

Job in Gresham, Multnomah County, Oregon, 97030, USA
Listing for: U.S. Bancorp
Full Time position
Listed on 2026-09-01
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Job Description & How to Apply Below
Position: Reliability Engineer 3 (Observability Specialist)

Reliability Observability Engineer 3

The Reliability Observability Engineer 3 is responsible for enabling reliable, measurable, and supportable application operations across a broad portfolio of production applications. This role helps ensure application teams have actionable visibility into service health, customer experience, dependencies, performance, and failure conditions through effective use of metrics, logs, traces, dashboards, alerts, synthetic monitoring, and service-level indicators.

As a senior-level Reliability Engineer specializing in observability, this role partners closely with product owners, application engineering teams, SRE teams, and business stakeholders to translate customer journeys and business outcomes into measurable reliability objectives. The role establishes and maintains best-practice processes for documenting, governing, reviewing, and improving user journeys, SLIs, SLOs, synthetic monitoring, dashboards, alerts, telemetry standards, and related observability assets.

The engineer provides senior technical guidance, identifies observability gaps through incident and performance analysis, drives continuous improvement, and helps ensure teams have the data, processes, and operating discipline needed to detect issues earlier, reduce customer impact, and improve overall service reliability.

Responsibilities include:

  • Leading the definition, documentation, implementation, and continuous improvement of Observability across Critical Customer Journeys, ensuring alignment between Observability Strategy, business outcomes, and reliability objectives.
  • Designing, implementing, and governing Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability metrics for enterprise applications and services.
  • Establishing and maintaining Observability Governance Frameworks for Dashboards, Alerts, Synthetic Monitoring, Telemetry Standards, and lifecycle management of observability assets.
  • Translating business and technical requirements into scalable Observability Architectures, including Instrumentation Standards, Monitoring Strategies, Tagging Frameworks, and Alerting Models.
  • Partnering with Product Owners, Application Engineering, Site Reliability Engineering (SRE), and Operations Teams to ensure applications are production-ready and fully instrumented for reliability measurement.
  • Developing and maintaining executive and operational Service Health Dashboards that provide insights into Availability, Latency, Customer Impact, Dependency Performance, and SLO Compliance.
  • Analyzing Telemetry Data, Incident Trends, Problem Records, and Alert Performance to identify observability gaps, reduce alert fatigue, and improve detection accuracy.
  • Providing technical leadership and mentorship on Distributed Tracing, Logging, Metrics Collection, Synthetic Testing, Application Performance Monitoring, Real User Monitoring, Monitoring Design Patterns, and Alert Governance Best Practices.
  • Leading the definition, documentation, and ongoing refinement of critical user journeys in partnership with product owners, engineering teams, SRE, operations, and business stakeholders to ensure observability practices are aligned to customer experience, business outcomes, and operational risk.
  • Defining, documenting, and governing appropriate service-level indicators and service-level objectives for applications and key capabilities, including availability, latency, error rate, throughput, dependency health, and other measurements that reflect meaningful customer and business impact.
  • Establishing and maintaining a best-practice process for identifying, approving, implementing, reviewing, and retiring user journeys, SLIs, SLOs, dashboards, monitors, synthetic tests, alerts, and related observability artifacts.
  • Translating product and engineering requirements into actionable observability designs that specify telemetry needs, measurement methods, tagging standards, dashboard requirements, alerting thresholds, ownership, evidence expectations, and operational runbook linkages.
  • Partnering with product and engineering teams during design, build, release, and production-readiness activities to ensure…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary