×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer, Data & Analytics

Job in Irvine, Orange County, California, 92713, USA
Listing for: Blizzard Entertainment
Full Time position
Listed on 2026-06-23
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 125000 - 150000 USD Yearly USD 125000.00 150000.00 YEAR
Job Description & How to Apply Below

Senior Site Reliability Engineer, Data & Analytics

The Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large‑scale data platforms, analytics pipelines, ML training pipelines, and inference services.

In addition to core SRE responsibilities, this role will build operational and automation tooling that reduces toil, speeds up issue resolution, and improves engineering velocity. This includes contributing to internal platform services such as shared tooling, data integrations, and access‑control patterns used across Blizzard.

The ideal candidate is a production‑minded SRE or platform engineer who is comfortable operating critical systems, writing software, and building tools that improve engineering efficiency without compromising reliability.

This role is open to candidates based in Irvine, CA or Albany, NY (hybrid or on‑site), as well as fully remote candidates.

Responsibilities
  • Participate in an on‑call rotation and drive incidents to resolution
  • Lead blameless postmortems and identify systemic reliability improvements
  • Partner with data, ML, and platform teams to improve batch, streaming, training, and inference workloads
  • Support ML training pipelines and inference services, including GPU workloads
  • Help define how data and ML services run on Kubernetes
  • Design and build automation and operational tooling (e.g., workflows, diagnostic tooling, runbooks) to reduce on‑call burden
  • Build and evolve centralized platform services, including shared tooling, data integrations, and access controls
  • Diagnose and resolve reliability, performance, and cost issues across distributed systems
  • Champion automation, documentation, and practices that reduce toil
  • Maintain infrastructure using Terraform and infrastructure‑as‑code principles
  • Improve CI/CD and Git Ops workflows (Jenkins, Git Hub Actions, ArgoCD)
  • Operate and improve containerized services on Kubernetes
  • Define and measure reliability using SLIs, SLOs, and error budgets
  • Run load tests, capacity modeling, and production validation
  • Build internal tools and paved paths that help teams operate safely and efficiently
Minimum Requirements
  • Experience operating reliable, distributed systems in SRE, platform, or similar roles
  • Experience with data, analytics, ML, or large‑scale distributed workloads
  • Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure
  • Experience building automation or internal tools (Python, Go, shell, etc.)
  • Experience with infrastructure‑as‑code (e.g., Terraform)
  • Experience with CI/CD or Git Ops systems (e.g., Jenkins, Git Hub Actions, ArgoCD)
  • Familiarity with observability (metrics, logs, traces, alerting, incident response)
  • Solid understanding of SRE concepts (SLIs, SLOs, error budgets, postmortems)
  • Experience using modern development and automation practices to improve reliability and efficiency
  • Experience building internal tooling, automation, or developer productivity systems
  • Strong communication skills with technical and cross‑functional partners
Bonus Points
  • Experience with data and ML systems (training pipelines, model serving, GPU workloads)
  • Experience with distributed systems and messaging (Kafka, Pub/Sub)
  • Experience working in Kubernetes‑based environments
  • Familiarity with observability tools (Prometheus, Grafana)
  • Experience operating systems in cloud environments (GCP, AWS)
Benefits
  • Medical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance
  • 401(k) with company match, tuition reimbursement, charitable donation matching
  • Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leave
  • Mental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, , rental insurance, and others
  • Relocation assistance if the company requires geographic mobility

In the U.S., the standard base pay range for this role is $ – $ annually. Compensation is based on experience, performance, and location.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, gender identity, age, marital status, veteran status, or disability status, among other characteristics.

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary