Senior/Site Reliability Engineer; SRE
Listed on 2026-08-29
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Site Reliability Engineer
We are looking for an experienced Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering with engineering teams to build systems that perform consistently under real-world demand. The ideal candidate brings deep production experience, a strong automation mindset, and a practical approach to incident response and continuous improvement.
Responsibilities include establishing measurable reliability standards for critical services by creating and maintaining service indicators, objectives, and error budget practices. The role also involves taking ownership of production stability by monitoring uptime, latency, and availability, and driving improvements that reduce operational risk. Additionally, the position requires leading live incident response efforts, coordinating troubleshooting during outages, and ensuring issues are resolved efficiently and thoroughly.
The Site Reliability Engineer will run blameless post-incident reviews, document findings clearly, and track corrective actions through completion. Design and enhance observability across logs, metrics, and distributed tracing using tools such as Datadog, Cloud Watch, Grafana, Open Telemetry, and Sentry. Improve alert quality and dashboard design so engineering teams can quickly identify meaningful system issues without unnecessary noise. Evaluate system behavior under load, uncover performance constraints, and recommend changes that improve scalability and resource efficiency.
Build automation and internal tooling that streamline operational work, strengthen deployment safety, and support incident management, debugging, and capacity planning. Contribute to infrastructure and delivery workflows across AWS, Terraform, Ansible, Linux, and Git Hub Actions with a focus on dependable releases and resilient systems. Partner with security and compliance stakeholders to support operational standards, audit readiness, and the integration of monitoring into broader engineering practices.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).