Senior Site Reliability Engineer
About The Role
Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation.
WhatYou'll Do
- System Architecture & Design:
Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems. Partner with development teams as a reliability consultant, reviewing designs and influencing architectural decisions to ensure new services are built with reliability, observability, and performance as core principles, not afterthoughts. - Automation & Software Development:
Write robust, performant, and maintainable code to automate operational tasks, and CI/CD pipelines. Build the internal tools, libraries, and frameworks that enable engineering teams to self-service their observability needs, reducing cognitive load and increasing their velocity. - Incident Response & Post-Mortem Analysis:
Participate in a 24/7 on‑call rotation, acting as a key technical leader and incident commander during critical service disruptions. Conduct deep, blameless root cause analyses (RCAs) that go beyond immediate fixes to identify and address systemic issues. Drive the implementation of corrective actions to prevent the recurrence of incidents. - Performance & Capacity Planning:
Proactively monitor, measure, and optimize system performance to ensure low latency and high efficiency. Gather and analyze metrics from operating systems and applications to assist in performance tuning and fault finding. Analyze usage patterns and historical data to forecast capacity needs, ensuring our platform stays ahead of customer demand.
- Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
- 5+ years of professional experience in a Site Reliability Engineering, Dev Ops, or Software Engineering role with a focus on infrastructure and operations.
- Strong programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript. You should be comfortable writing, testing, and deploying production-grade code.
- Deep knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53, Cloud Watch).
- Proven experience with Kubernetes in production (EKS preferred), including service exposure, networking, and availability engineering.
- A solid understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS, HTTP), and the architecture of modern distributed systems.
- Experience building and managing large‑scale monitoring and observability systems using tools like Datadog, Prometheus, Grafana, etc.
- Expertise in designing and implementing CI/CD pipelines using tools such as Git Hub Actions, ArgoCD, etc.
- Experience with distributed storage technologies (e.g., Amazon S3) and databases (e.g., Postgre
SQL, Scylla
DB, Click House, etc.). - Contributions to open‑source projects in the SRE, Dev Ops, or cloud‑native ecosystem.
Building the Future of Observability with AIResponsibilities
As a Senior SRE, you will be at the forefront of applying AI to solve our most critical reliability challenges. This is a hands‑on software development role where the “product” you build is an intelligent, automated reliability platform. Your responsibilities will include:
- Building AI‑Driven Automation:
Building and integrating solutions that leverage our AIOps platform. This involves writing the code that consumes signals from the AI system, correlates disparate data sources, automates responses to AI‑detected anomalies, and builds self‑healing systems triggered by predictive alerts. You will transform AI insights into concrete reliability improvements. - Leveraging AI for…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: