More jobs:
Site Reliability Engineer
Job in
San Ramon, Contra Costa County, California, 94583, USA
Listed on 2026-08-07
Listing for:
Prediktive
Full Time
position Listed on 2026-08-07
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Job Description & How to Apply Below
Site Reliability Engineer
We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California. Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.
Responsibilities- Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
- Automation and Toil Reduction:
Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning. - Incident Management:
Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly. - SLO and Error Budget Management:
Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability. - Monitoring and Observability:
Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency. - Post-Incident Reviews:
Conducting blameless post-mortems to find root causes and prevent repeat failures. - Capacity Planning:
Analyzing resource usage trends to forecast future infrastructure and scaling needs.
- Advanced level of English.
- 5+ years of experience working as a Site Reliability Engineer.
- 1+ years of experience within Microsoft Azure.
- Strong experience with Python, Bash or Go for automation purposes.
- Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
- Ability to keep a focused mind during high-severity production outages to lead teams effectively.
- Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
- Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
- Experience working with AI apps or agents.
- Bachelor's Degree in Computer Science, Systems Engineering or related fields.
- SaaS experience
- Long term positions
- Compensation in USD
- Paid time off
- Cool clients and products
- Work with great engineers
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×