×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer (SRE

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Methodic
Full Time position
Listed on 2026-08-22
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, Network Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE)

Open role

Site Reliability Engineer (SRE)

San Francisco, CA (On-site)

Responsibilities
  • Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers.
  • Perform regular capacity planning and load testing to ensure the platform can scale ahead of demand without performance degradation (validating that we can sustain extremely high transaction rates per tenant).
  • Improve deployment processes with strategies like blue-green or canary deployments to minimize risk and downtime during releases.
  • Collaborate with software engineers to design resilient architectures – for example, build redundancy and failover capabilities into critical services (so that even if one component fails, the system remains operational).
  • Participate in on-call rotations to respond to and resolve production incidents; lead blameless post-mortems to identify root causes and implement corrective actions.
  • Automate routine operational tasks (from simple scripts to more complex tooling) to reduce manual work and error potential.
  • Document reliability-related procedures and best practices, ensuring knowledge is shared and systems are well-understood by the team.
Requirements
  • 5+ years in a Site Reliability Engineering or similar role.
  • Strong coding/scripting abilities (Python, Go, or other) for building automation and tooling.
  • Deep knowledge of systems monitoring and observability — experience with tools like Grafana, Datadog, Prometheus, etc., and the ability to interpret system metrics to spot problems.
  • Understanding of high-availability design and distributed systems principles (load balancing, consensus, graceful degradation, etc.).
  • Experience with incident management and a track record of improving systems based on lessons learned.
  • Performance tuning experience — ability to use profiling and stress testing tools to find and fix bottlenecks.
  • A collaborative mindset, capable of working with development teams to ensure reliability is built in from the start, not just after issues occur.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary