Senior Site Reliability Engineer
Listed on 2026-09-10
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Cybersecurity
Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend oncall, the Site Reliability team keeps the Salesforce cloud and our customers protected.
TheExperience
As an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI-powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability leveraging cutting-edge software engineering practices within SRE function and AI-driven insights, you will help transform how services are built, monitored, and operated — ensuring that Salesforce delivers always-on, high-performance experiences to customers worldwide.
Build and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient — while continuously improving the balance between operational work and engineering innovation.
Reliability as the Priority :
Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.
Engineering for Operations :
Apply software engineering practices — automation, monitoring, self-healing systems — to eliminate toil and improve operational efficiency.
Incident Management :
Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.
Continuous Improvement :
Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).
Collaboration with Development :
Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.
Long-Term Focus :
Leverage AI-driven automation to eliminate manual workflows, enabling the team to focus on complex problem-solving and strategic innovation while reducing operational overhead to less than 20% of capacity.
Lead incident detection, response, and resolution—driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.
Lead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak performance and reliability.
Understanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).
Independently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).
Architect and build production-grade observability solutions — monitoring, logging, alerting, and tracing systems — that enable proactive detection and autonomous remediation.
Design and implement AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents (MCP-based).
Drive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning.
Ensuring that work…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).