Sr. Staff Production Engineer - Data Platform
Listed on 2026-08-22
-
Software Development
Overview
Databricks is building the world’s best data and AI infrastructure platform, enabling data teams to solve the toughest problems—from transportation innovations to accelerating medical breakthroughs. Founded by engineers and customer‑obsessed, Databricks tackles technical challenges across a massive, multi‑cloud (AWS, Azure, GCP) stack that powers thousands of demanding workloads.
RoleAs a Senior Staff Production Engineer you will lead the strategic vision for the operational stability of our internal "Databricks‑on‑Databricks" environment. You will transition infrastructure from traditional SRE models toward an agent‑driven, self‑healing architecture, ensuring that the platform—and the agents operating within it—remain rock‑solid for mission‑critical customer workloads.
Impact You Will Have- Architect Agentic Reliability: define and drive the design of future "self‑healing" infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers.
- Data Platform Optimization: own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for compute, storage, and control plane services used by thousands of Databricks engineers.
- High‑Scale Operational Excellence: establish the next generation of "Change Safety" protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions.
- Leadership in Chaos & Scale: serve as a technical bar‑raiser, evangelizing modern SRE practices—including Chaos Engineering—to navigate the structural transformation toward agentic, autonomous systems.
- BS/MS/PhD in Computer Science or a related field.
- 10+ years of production‑level experience as a Software Engineer or SRE in highly distributed, multi‑cloud environments.
- Engineering Persona: write code to solve operational problems; build frameworks, automation, and tooling (Scala, Java, Go, or Python) to eliminate toil.
- Platform & AI Mindset: deep understanding of distributed data platforms and a passion for leveraging AI/ML to revolutionize infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus.
- Operational Grit: proven ability to remain calm and decisive under pressure, navigate large‑scale distributed systems, and drive incident‑to‑roadmap loops.
- Strategic Influence: experience building long‑range technical roadmaps and driving cross‑functional alignment, comfortable challenging senior leadership with data‑driven insights.
Zone 1 Pay Range: $228,600 — $314,250 USD.
EEO StatementIndividuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio‑economic status, veteran status, and other protected characteristics.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).