Senior Production Engineer - SRE
Listed on 2026-09-17
-
IT/Tech
SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations
What You Can Expect
In this role, you will be responsible for the reliability, scalability, and operational excellence of Zoom's government and military-facing products. A core focus is deploying and evolving our observability platform, working across Kubernetes, Terraform, and cloud infrastructure to build the systems and frameworks that set the standard for how the team operates. This is a hands-on, on-call engineering role: you will own the systems you build, respond to production incidents with urgency, and deliver permanent improvements.
Aboutthe Team
We build and own the cloud infrastructure behind Zoom's government products, operating with full end-to-end accountability. Our team moves fast, solves hard problems, and makes a direct impact where reliability is non-negotiable.
ResponsibilitiesArchitecting and owning the end-to-end deployment lifecycle of Zoom's observability platform across AWS Gov Cloud and Oracle OCI, streamlining processes, eliminating manual steps, and setting the standard for how teams deploy at scale
Designing and building Terraform modules and providers from scratch, alongside Git Lab CI/CD pipelines, to automate infrastructure provisioning and configuration management across development and production environments
Operating and improving production Kubernetes workloads at scale, leading incident response, driving root cause analysis, and delivering permanent fixes that reduce outage risk and improve system resilience
Developing reusable frameworks, runbooks, and operational standards that engineering teams across the organization can adopt — raising the bar on observability onboarding, alerting coverage, and operational readiness
Collaborating with engineering teams to define and track SLOs and SLIs, contribute to architecture and design reviews, and systematically eliminate toil through automation and thoughtful infrastructure design
Holds U.S. Citizenship or Lawful Permanent Resident (Green Card) status
Brings 4–5+ years of hands-on experience in Platform Engineering, Site Reliability Engineering, Dev Ops, or Production Engineering in a live production environment
Demonstrates production-grade Kubernetes experience, including workload operations, troubleshooting, and cluster-level architectural design at scale
Authors Terraform modules and providers from scratch, not limited to applying or executing existing configurations
Owns incidents end-to-end - proven on-call experience including structured triage, resolution, and post-incident review with documented root cause analysis
Deploys, operates, or improves observability or monitoring platforms at production scale, with scripting proficiency in Python and/or Bash for automation and operational tooling
Communicates clearly in writing - produces runbooks, incident reports, methods of procedure, and architectural documentation to a high standard
Has hands-on Datadog experience building observability frameworks in a production environment, ideally within a government, FedRAMP, or high-compliance cloud context
Experience operating within FedRAMP, AWS Gov Cloud, or other regulated or compliance-driven cloud environments
Hands-on experience with popular open-source observability tooling in a production context
Background supporting 24/7 or mission-critical production operations
Experience driving cross-team adoption of operational standards or tooling frameworks
Oracle OCI hands-on operational experience
Familiarity with container image security and CVE remediation workflows
Minimum:
$98,900.00
Maximum:
$
In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.
N…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).