×
Register Here to Apply for Jobs or Post Jobs. X

Senior Manager, Site Reliability Engineering

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: Intuit Inc.
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 222000 - 300500 USD Yearly USD 222000.00 300500.00 YEAR
Job Description & How to Apply Below

About the Team

Intuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps Quick Books, Turbo Tax, Credit Karma, and Mailchimp running for hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency tooling, and incident response capability that underpins Intuit's money-movement and fintech services — where availability, data integrity, and trust are non-negotiable.

The

Opportunity

We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10–15 systems and reliability engineers responsible for the availability, performance, and operational health of Fintech Platform services running in AWS. This leader owns the strategy and execution behind operational excellence: driving toward a 99.999% availability bar, maturing incident management practices, and building self-healing, well-instrumented infrastructure at scale.

This is a player-coach role. You will set technical direction and organizational strategy while staying close to the systems — reviewing designs, joining incident bridges, and coaching engineers through complex production issues. You'll partner closely with software engineering, product, security, and other SRE/infrastructure leaders across Intuit to raise the bar on reliability company-wide.

A defining priority for this role is AI Ops: embedding AI-driven, autonomous operations into how the team runs infrastructure. You will lead the shift from manual, human-triggered response toward self-healing systems that detect, diagnose, and remediate issues autonomously — reducing developer toil, cutting MTTR, and freeing engineering capacity to focus on higher-value work. Done well, this delivers 3x the operational impact of the team today and directly accelerates the pace at which we deliver value to customers.

Responsibilities

Responsibilities
  • Own end-to-end operational excellence for Fintech Platform services: define and drive the strategy for achieving and sustaining 99.999% availability across customer-facing and internal systems.

  • Lead, grow, and directly manage a team of 10–15 systems/site reliability engineers — hiring, mentoring, setting goals, and developing the next generation of technical leaders.

  • Act as a hands‑on technical leader: participate in architecture and design reviews, write and review code/IaC where needed, and dive into production systems alongside the team.

  • AI Ops:
    Driving 3x Impact Through Autonomous Operations-

    Define and execute an AI Ops roadmap that embeds autonomous detection, diagnosis, and remediation into production systems, targeting a 3x improvement

  • Identify high-toil, repetitive operational workflows and systematically replace them with autonomous agents and automation, freeing engineers to focus on higher‑leverage engineering work.

  • Measure and report on toil reduction, automation coverage, and velocity gains, tying AI Ops investment directly to faster, safer delivery of customer value.

  • Drive incident management maturity — own the incident command process, lead or oversee response for high‑severity (P1/P2) incidents, and ensure rigorous root‑cause analysis and blameless postmortems.

  • Build and scale AWS cloud infrastructure (compute, networking, storage, container orchestration) with a focus on resiliency, auto‑remediation, chaos engineering, and multi‑AZ/multi‑region failover.

  • Define and report on SLOs/SLIs, error budgets, and availability metrics; use data to prioritize reliability investments and reduce toil through automation.

  • Partner with software engineering, product management, security, and compliance teams to embed reliability, observability, and operational readiness into the software development…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary