More jobs:
Site Reliability Engineer (SRE
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-08-22
Listing for:
Methodic
Full Time
position Listed on 2026-08-22
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, Network Engineer
Job Description & How to Apply Below
Open role
Site Reliability Engineer (SRE)
San Francisco, CA (On-site)
Responsibilities- Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers.
- Perform regular capacity planning and load testing to ensure the platform can scale ahead of demand without performance degradation (validating that we can sustain extremely high transaction rates per tenant).
- Improve deployment processes with strategies like blue-green or canary deployments to minimize risk and downtime during releases.
- Collaborate with software engineers to design resilient architectures – for example, build redundancy and failover capabilities into critical services (so that even if one component fails, the system remains operational).
- Participate in on-call rotations to respond to and resolve production incidents; lead blameless post-mortems to identify root causes and implement corrective actions.
- Automate routine operational tasks (from simple scripts to more complex tooling) to reduce manual work and error potential.
- Document reliability-related procedures and best practices, ensuring knowledge is shared and systems are well-understood by the team.
- 5+ years in a Site Reliability Engineering or similar role.
- Strong coding/scripting abilities (Python, Go, or other) for building automation and tooling.
- Deep knowledge of systems monitoring and observability — experience with tools like Grafana, Datadog, Prometheus, etc., and the ability to interpret system metrics to spot problems.
- Understanding of high-availability design and distributed systems principles (load balancing, consensus, graceful degradation, etc.).
- Experience with incident management and a track record of improving systems based on lessons learned.
- Performance tuning experience — ability to use profiling and stress testing tools to find and fix bottlenecks.
- A collaborative mindset, capable of working with development teams to ensure reliability is built in from the start, not just after issues occur.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×