×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Sycamore
Full Time position
Listed on 2026-08-13
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below

Own how the platform behaves over time: service levels, observability, capacity, and what happens when something breaks.

About Sycamore

Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate AI agents that take on real operational work, with the security and control large organizations need.

We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.

The role

You will build the infrastructure and control-plane software that lets Sycamore’s services, mutable development sandboxes, and immutable deployed applications run securely from development through production.

This is software engineering for consequential distributed systems, not cloud administration. You will write control-plane services and Kubernetes controllers, codify infrastructure and delivery paths, design identity and network boundaries, operate tenant-isolated data systems, and debug failures that cross application, cluster, database, and customer‑network layers.

You own more than whether a cluster is healthy or a deployment completed. You own whether the system is isolated, observable, recoverable, cost‑aware, and safe to change, and whether the next deployment is more repeatable than the last.

Infrastructure is the foundation beneath Product and Core AI. Errors here can affect every application and customer, so we maintain an especially high bar for systems judgment, production ownership, and recovery discipline.

Working on Site Reliability

Cloud Infrastructure owns the shape of the system. This role owns its behavior over time.

That means service level objectives and the error budgets that make them mean something, an observability estate managed as code rather than clicked together, the promotion gates that decide whether a release reaches production, capacity and cold‑start performance, and incident response from page through postmortem to the structural fix.

We are honest about where this stands. We have a large monitor fleet and a metric taxonomy that is genuinely enforced, with policy checks that reject a monitor lacking a runbook link or proper scoping. What we mostly do not have yet is the practice on top: few service level objectives, no error budgets, no burn‑rate alerting, and a production promotion gate that currently warns rather than blocks.

Building that practice is the role, not a side project within it.

The reliability problems themselves are specific. Streaming endpoints that hold a database connection for the length of an upstream response. Cold starts on preemptible capacity racing a startup probe. Memory limits that censor the very measurement needed to size them. Warm pools that silently fall back to cold provisioning. These are measured, documented, and open.

Our incident tooling is also unusual: alert triage is substantially agent‑driven, with automated investigation and ticket filing, and a decay process that nominates monitors nobody has acted on for deletion. You would own and extend that.

What the work looks like

In one week, you might:

  • Trace a failed application deployment across artifact creation, image building, registry publication, admission controls, routing, workload identity, health checks, and rollback.
  • Build or extend a control‑plane service or Kubernetes controller that provisions and reconciles tenant or sandbox resources safely.
  • Design an edge‑authentication, service‑identity, or delegated‑access flow and verify that untrusted workloads cannot cross the boundary.
  • Investigate latency or reliability across gateways, containers, managed databases, object storage, cluster scheduling, and customer DNS or networks.
  • Review an infrastructure change for blast radius, execute a staged rollout, watch health signals, and prove the rollback path.
  • Restore tenant data to a known point, verify the result, and turn the exercise into a repeatable recovery procedure.
  • Improve sandbox startup time, deployment throughput, resource efficiency, or capacity without weakening isolation.
  • Work with a customer’s infrastructure or…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary