Site Reliability Engineer
Listed on 2026-09-09
-
Software Development
AI Engineer (Applied/Software)
About the team
Sight Machine is built on the shoulders of a unique, robust and highly scalable Infrastructure as Code model. This enables the creation and operation of customer instances in our ecosystem in a standardized and simplified manner. We are looking for team members to help us build, maintain, and improve the infrastructure that makes Sight Machine the leading provider of Manufacturing Data Pipelines and Analytics.
Great things happen when people can bring their authentic selves to work. We empower all of our team members to share their perspectives, passions and experiences because collectively we make a better, stronger team through always “open communications” mind.
Our team collaborates closely with peers & cross functional stakeholders throughout the business, our clients on the forefront of digital transformation, and the cutting edge of digital manufacturing thought leadership.
Sight Machine has offices in San Francisco, CA and Ann Arbor, Mi. We do have a remote-friendly culture with people based all around the US and the rest of the world. For this role in particular, the ideal candidate is located near either of our offices and willing to work in a hybrid capacity. We would still consider 100% remote for exceptional candidates if they aren’t located near an office.
Aboutthe role
Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight Machine's platform. You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM gateways, agent orchestration, and the operational patterns that come with non-deterministic workloads.
This is a senior level IC role. You'll help set and drive technical direction for infrastructure and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems problems while still being hands-on with code, infrastructure, and incidents.
Success requires deep technical range, sound judgment on risk vs. customer impact, and the ability to influence architecture decisions across Development Engineering without formal authority.
What You’ll Actually Work On- Champion an agentic-AI-first engineering mindset: identify where AI-driven automation and agent-based tooling can replace manual toil, and hold that work to the same quality, testing, and reliability bar as any other production system
- Evolve reliability practices for meeting reliability SLO’s, error budgets, drive incident postmortems to systemic (not just symptomatic) fixes, and lead reliability reviews for new services before they hit production
- Troubleshoot and resolve the org's most complex, cross-layer systems problems CI/CD, container orchestration, networking, OS, cloud resources, databases, and increasingly, agentic AI/LLM orchestration layers
- Design, build, and operate the infrastructure supporting agentic AI workloads, LLM gateway routing, agent orchestration frameworks, monitoring of non-deterministic/AI-driven services, and the operational tooling needed to run them reliably at scale
- Architect and instrument monitoring, alerting, and observability infrastructure for critical services, with an eye toward what "critical" means for AI-driven systems specifically
- Author and continuously improve operational runbooks and automation, increasingly incorporating agentic/AI-assisted tooling (e.g., automated triage, AI-assisted incident response) where it measurably reduces toil
- Design and build internal platforms and developer tooling that other engineers build on top of
- Participate in on-call coverage and help evolve the program as we scale including escalation paths and reducing avoidable pages through better automation
- Bring a startup mindset of daily engagement: staying close to what's breaking, what customers are hitting, and where the team needs help, even outside a formal ticket or rotation
- Mentor senior and mid-level engineers; act as a technical sounding board across teams
- Proactively identify and drive cross-team initiatives that improve…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).