Senior DevOps Engineer
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability, Systems Engineer, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Zof AI is seeking a Senior Dev Ops Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane:
Kubernetes and container orchestration, CI/CD pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large agent fleets affordable. If you have worked as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or Infrastructure Engineer, this is that discipline at Zof AI. The ideal candidate has operated production infrastructure at scale and treats security, reliability, and cost per agent run as constraints they personally own.
Engineering
· Mid to Senior
· Full-time
· On-site
· San Francisco, CA
- Design and operate the sandboxed environments where agents reproduce defects and validate fixes.
- Own Kubernetes, container, and compute infrastructure end to end.
- Build CI/CD pipelines that let engineers ship safely many times a day.
- Harden isolation boundaries so untrusted customer code stays inside its sandbox.
- Instrument fleet health with metrics, logs, tracing, and alerting that catch failures early.
- Drive down cost per agent run through scheduling, autoscaling, and capacity work.
- Automate provisioning, deployment, rollback, and environment management.
- Partner with engineers to make infrastructure fast and safe to build on.
- Experience running production infrastructure on a major cloud platform.
- Working knowledge of Kubernetes, containers, and orchestration.
- Experience building CI/CD pipelines, deployment automation, and infrastructure as code.
- Familiarity with observability, monitoring, and on-call practice.
- Judgment about security, reliability, and cost trade-offs.
- Daily use of AI tools to automate operational and engineering work.
- Clear written and verbal communication.
- Comfort operating in a fast-moving environment.
- Experience with sandboxing or multi-tenant isolation tooling such as gVisor or Firecracker.
- Experience with Terraform, Pulumi, or similar infrastructure as code tooling.
- Experience running large batch or job-based workloads cost efficiently.
- Experience in early-stage infrastructure or platform teams.
Must be able to run infrastructure for large agent workloads and use AI tools to automate operational work
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).