×
Register Here to Apply for Jobs or Post Jobs. X

Senior Data Ops Engineer, Data Activation & Products - Activision

Job in Santa Monica, Los Angeles County, California, 90403, USA
Listing for: ACTIVISION PUBLISHING, INC.
Full Time position
Listed on 2026-09-22
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 102800 - 190204 USD Yearly USD 102800.00 190204.00 YEAR
Job Description & How to Apply Below

Job Title:

Senior Data Ops Engineer, Data Activation & Products - Activision Requisition : R028025

Mission

We are looking for a Site Reliability Engineer to help improve the reliability, observability, and operational maturity of our data platforms, Kubernetes-based deployment systems, internal applications, and cloud environments. This role sits at the intersection of SRE, Dev Ops, and data. The ideal candidate is comfortable operating production systems, troubleshooting across infrastructure and applications, and helping teams deploy and support services more safely.

The role does not require someone to be a data engineer, but they should be excited about the systems that support modern data engineering, including Databricks, Spark, Airflow/Astronomer, streaming pipelines, event systems, internal tools, and backend services. We are especially interested in a creative, curious engineer who enjoys learning new technology and using AI-accelerated development practices to solve problems faster and more thoughtfully.

You should be excited to experiment with agentic development tools, automation frameworks, and emerging platform capabilities, while applying sound engineering judgment.

Responsibilities
  • Monitor service health, respond to alerts, and participate in incident response for cloud, Kubernetes, application, and data platform environments.
  • Investigate reliability issues across Kubernetes, networking, DNS, application runtime behavior, Databricks jobs, Spark workloads, orchestration systems, event systems, and dependent services.
  • Support the reliability of internal applications, APIs, workers, streaming consumers, event‑driven services, and deployment workflows used by data engineering and business teams.
  • Build and maintain dashboards, alerting, runbooks, and operational documentation that improve detection and recovery speed.
  • Improve observability for Databricks environments, including job health, Spark streaming workloads, structured streaming metrics, cluster behavior, failures, latency, throughput, and cost signals.
  • Help route Spark streaming metrics, operational logs, event‑system signals, and platform health signals into monitoring tools such as Grafana.
  • Contribute to alerting patterns for Databricks workflows, Airflow/Astronomer DAGs, dbt jobs, data freshness, pipeline failures, event lag, dead‑letter queues, and production data dependencies.
  • Contribute scripts and automation that reduce repetitive operational work and improve environment hygiene.
  • Support release and deployment reliability by validating changes, improving rollback readiness, and strengthening change safety.
  • Partner with data engineers, analytics engineers, and software engineers to improve reliability across pipelines, services, internal tools, event systems, and data products.
  • Participate in post‑incident follow‑up and help close corrective actions that prevent recurrence.
  • Support platform modernization and migration efforts, including orchestration platform changes, deployment system improvements, and shared reliability standards.
Qualifications
  • 5+ years of experience in SRE, Dev Ops, cloud infrastructure, platform engineering, software engineering, data platform operations, or related production‑support roles.
  • Hands‑on experience supporting Kubernetes‑based workloads, deployment systems, cloud infrastructure, or production application environments.
  • Familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git‑based workflows, and scripting in Bash, Python, or similar languages.
  • Experience with monitoring, logs, metrics, dashboards, alerting, and incident management practices.
  • Comfort working with event systems such as Kafka, Google Pub/Sub, Kinesis, or similar technologies, including topics, subscriptions, consumers,…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary