×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Instrumental Inc.
Full Time position
Listed on 2026-09-13
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 165000 USD Yearly USD 140000.00 165000.00 YEAR
Job Description & How to Apply Below

Instrumental builds the manufacturing acceleration platform behind the world’s most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, Cisco, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production.

The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface—we believe we must have both the best technology and the best access to that technology to win.

As a Site Reliability Engineer
, you’ll operate, improve, and scale our AWS-based SaaS platform. You’ll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational excellence. You’ll participate in a bi-weekly on-call rotation, but the goal isn’t simply to keep systems running—it’s to continuously engineer away the operational complexity that comes with scaling our platform and customer base.

Requirements
  • 3–4 years of experience in Site Reliability Engineering, Dev Ops, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments.
  • Strong hands‑on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3.
  • Experience managing infrastructure using Terraform or other Infrastructure as Code technologies.
  • Experience designing and supporting CI/CD pipelines using Git Hub Actions, Jenkins, Git Lab CI/CD, or similar platforms.
  • Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM.
  • Experience with Docker and Kubernetes.
  • Scripting experience with Python and/or Bash.
  • Experience supporting production environments through an on‑call rotation, including incident response and root cause analysis.
  • Proven ability to take ownership of production issues and drive them through investigation, remediation, and long‑term resolution.
Who You Are
  • Dead serious about performance, scalability, and reliability (PSR):
    You care deeply about how systems behave in the real world and continuously look for ways to make them more reliable, scalable, observable, and supportable.
  • Automation, automation, automation:
    If something is repetitive, manual, or error‑prone, your first instinct is to automate it and make it disappear.
  • An engineer at heart:
    You don’t want to repeatedly fight the same fires. You look for the underlying cause and build durable engineering solutions that reduce operational toil and technical debt.
  • Strong systems thinker:
    You understand how infrastructure, applications, networks, deployments, monitoring, and people interact—and can troubleshoot complex production issues across those boundaries.
  • Collaborative and reliable:
    You partner closely with software engineers to make services production‑ready, improve operational workflows, and build reliability into systems before they become problems.
  • Comfortable with growth and ambiguity:
    You’re comfortable making good decisions without perfect information and adapting as the platform, customer base, and company scale quickly.
Nice To Have
  • Experience working in a high‑growth B2B SaaS environment.
  • Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
  • Experience building internal tooling and automation to eliminate operational…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary