×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, Ads

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Jackalope Digital LLC
Full Time position
Listed on 2026-08-31
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 217000 - 304000 USD Yearly USD 217000.00 304000.00 YEAR
Job Description & How to Apply Below
Position: Staff Site Reliability Engineer, Ads

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.

For more information, visit

This role is remote friendly. Reddit has a flexible first workforce.

The Ads organization powers Reddit's advertising platform, enabling advertisers to reach highly engaged communities while helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser success, revenue generation, and user experience.

The Ads Reliability team partners closely with Ads Engineering teams to improve reliability, scalability, operational excellence, and developer productivity across Reddit's advertising ecosystem.

We're looking for a Staff Site Reliability Engineer who will define and provide technical leadership for reliability initiatives across the Ads organization and help shape the future of Ads infrastructure at Reddit.

What you’ll do:
  • Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Partner with engineering leadership to develop a roadmap to improve reliability, scalability, operational excellence, and engineering efficiency across the Ads organization.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Drive architecture reviews and influence technical decisions impacting critical revenue-generating systems.
  • Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.
  • Identify systemic reliability risks and drive long-term solutions that improve platform resilience.
  • Establish reliability metrics around advertiser‑critical user journeys such as campaign creation, ad delivery, auction participation, reporting, attribution, and billing.
  • Mentor engineers and provide technical leadership across multiple teams.
  • Influence roadmap planning and ensure reliability considerations are incorporated into product and infrastructure investments.
What We’re Looking For
  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large scale distributed systems.
  • Strong experience evolving high traffic, user‑facing production environments.
  • Strong cross‑functional collaborations skills to lead and influence projects driving operational excellence.
  • Deep expertise in modern distributed systems, scale engineering, and cloud‑native architectures.
  • Experience designing highly‑available systems with strong operational and reliability practices.
  • Strong software engineering skills in languages like general‑purpose backend languages like Go.
  • Strong understanding of observability systems including metrics, logging, tracing, and alerting.
  • Experience improving reliability through SLOs, automation, incident management, and performance optimization.
  • Demonstrated ability to troubleshoot complex issues across a modern distributed system stack.
  • Strong collaboration and communication skills with the ability to influence technical direction across teams.
Nice to Have
  • Experience supporting advertising technology platforms or other large‑scale revenue‑critical systems.
  • Deep understanding of reliability challenges associated with ad‑serving, real‑time auctions, budget pacing, campaign delivery, measurement, attribution, or billing systems.
  • Experience operating high‑QPS, low‑latency services where latency directly impacts business outcomes.
  • Experience establishing reliability programs that deliver meaningful, measurable business outcomes
  • Experience with Kubernetes, cloud infrastructure, and large‑scale distributed systems.
  • Familiarity with Kafka, Click House, Spark, Flink, Big Query, or similar large‑scale data platforms.
  • Experience partnering with Product, Data Science, and Ads Engineering organizations.
  • Experience supporting…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary