×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in San Ramon, Contra Costa County, California, 94583, USA
Listing for: Prediktive
Full Time position
Listed on 2026-08-07
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Job Description & How to Apply Below

Site Reliability Engineer

We are looking for a Site Reliability Engineer based in Latin America to work on a long-term project for one of our clients, a software company based in San Ramon, California. Our client provides cloud solutions trusted by governments worldwide to accelerate their digital transformation, deliver vital services, and build stronger communities.

Responsibilities
  • Lead deep technical analysis during critical incidents, coordinating cross-functional teams to identify root causes, restore services quickly, and drive long-term reliability improvements.
  • Automation and Toil Reduction:
    Writing code and scripts (Python, Go, or Bash) to eliminate repetitive manual work and handle provisioning.
  • Incident Management:
    Responding to system alerts, troubleshooting production issues, joining on-call rotations, and restoring services quickly.
  • SLO and Error Budget Management:
    Establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance feature delivery speed with system stability.
  • Monitoring and Observability:
    Building dashboards and setting up tools (such as Prometheus, Grafana, or Datadog) to track system health and latency.
  • Post-Incident Reviews:
    Conducting blameless post-mortems to find root causes and prevent repeat failures.
  • Capacity Planning:
    Analyzing resource usage trends to forecast future infrastructure and scaling needs.
Requirements
  • Advanced level of English.
  • 5+ years of experience working as a Site Reliability Engineer.
  • 1+ years of experience within Microsoft Azure.
  • Strong experience with Python, Bash or Go for automation purposes.
  • Experience working with Monitoring and Observability tools such as Prometheus, Grafana, or Datadog.
  • Ability to keep a focused mind during high-severity production outages to lead teams effectively.
  • Proven experience explaining complex infrastructure failures to non-technical business stakeholders in clear, simple terms.
  • Strong prioritization instincts, with the ability to quickly determine whether to resolve a live issue manually or engineer a permanent automated solution.
  • Experience working with AI apps or agents.
Bonus Points
  • Bachelor's Degree in Computer Science, Systems Engineering or related fields.
  • SaaS experience
What We Offer
  • Long term positions
  • Compensation in USD
  • Paid time off
  • Cool clients and products
  • Work with great engineers
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary