×
Register Here to Apply for Jobs or Post Jobs. X

Observability SRE Manager, Apple Services Engineering

Job in Seattle, King County, Washington, 98127, USA
Listing for: Apple Inc.
Full Time position
Listed on 2026-08-31
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Project Manager
Salary/Wage Range or Industry Benchmark: 225600 - 338400 USD Yearly USD 225600.00 338400.00 YEAR
Job Description & How to Apply Below

Observability SRE Manager, Apple Services Engineering

Seattle, Washington, United States Software and Services

People at Apple don't just build products, they craft the kind of experiences that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here!

Join Apple, and help us leave the world better than we found it.

The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love.

Apple's observability and monitoring platforms are the nervous system behind the reliability of Apple's cloud services, giving thousands of engineers the visibility they need to detect, diagnose, and resolve issues before they impact customers. Our cloud monitoring platform analyzes billions of metrics per minute and is the first place incident responders turn when something goes wrong, regardless of system scale or complexity.

We're looking for a senior SRE leader to own and evolve this platform: the metrics, logging, tracing, and alerting infrastructure that underpins operational excellence across Apple Services Engineering.

Description

You'll set technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software, infrastructure, finance, and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple scale.

This is a senior leadership position that defines where our observability platform reliability, scalability and performance goes next, how our SRE practice evolves, including how AI reshapes it, and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter, integrating monitoring seamlessly across disparate infrastructures, hardware, software, application, and network layers, at a scale built to reach every user on the planet.

You'll build a team culture that makes SRE sustainable, rewarding, and central to how Apple ships services.

The successful candidate has a strong aptitude for both technical leadership and people management, with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams, driving complex cross-functional initiatives, and thriving under pressure, while creating an inclusive, high-trust team culture where engineers do their best work.

We believe AI will fundamentally reshape how SRE is practiced, from anomaly detection and root-cause analysis to capacity planning and toil elimination, and we're looking for a leader who shares that conviction and can drive that transformation across the organization.

Responsibilities
  • Technical & Operational Leadership
  • Own the reliability, availability, and performance of Apple's observability platform (metrics, logging, tracing, alerting) and the self-service capabilities built on top of it
  • Lead staging and production environments for the observability platform with the goal of maximizing availability
  • Ensure the platform accurately monitors the health of every application and piece of infrastructure across the Apple ecosystem, the "central nervous system" that other engineering teams and incident responders reach for first
  • Define and drive the strategic roadmap for observability infrastructure in partnership with SRE, engineering, and product stakeholders
  • Establish and refine SRE practices including SLOs, error budgets, capacity planning, scale testing, disaster recovery, and change management
  • Guide deep dives into systemic and latent reliability issues spanning the full stack (hardware, software, application, and network), partnering with software and systems engineers to drive fixes to resolution
  • Drive incident response, post-incident reviews, and systemic improvements that reduce operational toil
  • Champion automation to eliminate manual processes through tooling, self-service platforms, and APIs…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary