Staff Site Reliability Engineer-Production Operations
Listed on 2026-08-03
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
About Us
Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity solutions, we’re addressing the challenges of electric vehicles through technology that will set the standards for software-defined vehicles around the world.
Role SummaryWe are looking for an SRE Lead to serve as the senior technical leader and player-coach for Prod Ops. This is a hybrid role: you will set the technical direction of the team and lead from the front during incidents, while also growing and managing a small group of exceptional engineers as the function scales.
Responsibilities Lead incident coordination and communicationDrive incident progress and coordinate cross-functional response as the central nervous system during a crisis. Own executive, customer, and parent-company communications, providing production expertise and communication leadership so engineers can concentrate on the technical problem. Maintain an accurate, high-level model of the full production ecosystem spanning Cloud, vehicle development pipelines, and Product Security.
Facilitate modern, blameless post-incident learningRun post-incident reviews using Learning From Incidents (LFI) principles and HOWIE-style reporting. Move the organization away from the search for a single root cause and toward understanding how tooling, context, and multiple latent conditions combined to produce failure. Surface weaknesses in observability, process, testing, and tooling that impaired our ability to detect, mitigate, and recover.
Own and prioritize systemic action itemsManage the backlog of action items generated by reviews. While Prod Ops does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact.
Measure and verify efficacyOnce the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate.
Build the automation and observability backboneDesign and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team’s impact.
Lead and grow the teamSet technical standards and operating rhythm for Prod Ops. Mentor and develop engineers, and as the team scales, take on hiring and people management while preserving the minimal-headcount, maximum-automation philosophy.
Qualifications- 6+ years of experience in SRE, production/platform engineering, or systems engineering for large-scale distributed systems, including senior technical leadership or lead responsibilities.
- Proven incident command experience: you have coordinated high-severity, cross-functional incidents and led communication with executives and external stakeholders under pressure.
- Strong systems engineering fundamentals, including distributed systems, networking, cloud infrastructure, and an instinct for how complex systems fail.
- Hands-on coding ability (e.g., Python, Go, or similar) sufficient to build automation, tooling, and integrations. This is not a code-free management role.
- Deep observability expertise: instrumentation, metrics, logging, tracing, alerting, dashboards, and SLO/SLI design (Datadog or comparable platforms).
- Fluency in modern reliability and post-incident practice: blameless reviews, Learning From Incidents (LFI), HOWIE, and systemic (non-single-root-cause) analysis.
- A demonstrated bias toward eliminating toil through automation, and enthusiasm for using AI agents as force multipliers.
- Excellent written and verbal communication; able to translate technical detail for executive and cross-company audiences.
- Experience mentoring engineers, with the judgment and appetite to grow into formal people management.
- Experience across both cloud services and hardware/vehicle or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).