Principal DevOps Engineer
San Jose, Santa Clara County, California, 95199, USA
Listed on 2026-09-12
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
We are seeking a Principal Dev Ops Engineer who combines deep technical expertise with broad system understanding. This engineer should be capable of diving into a wide range of services and identifying systemic issues across architecture, CI/CD flow, and containerization environments. This role requires technical leadership, analytical skill, and cross-team collaboration to drive reliability, scalability, and modernization.
About the TeamAt Zoom, we’re building the next generation of Cloud and Colocation (Colo) infrastructure that powers seamless communication and collaboration for millions of users worldwide.
Responsibilities- Leading deep-dive investigations across diverse services and environments. Working on real time media systems to web, team chat and AI to uncover architectural or operational bottlenecks.
- Designing and implementing improvements in deployment pipelines, orchestration frameworks, andCI/CD automation to increase reliability and release velocity.
- Working closely with product and service owners to enhance containerization strategy, improve resource efficiency, and reduce operational friction.
- Partnering with the Meeting Dev Ops and Cloud Infra teams to modernize hybrid infrastructures panning colocation data centers, AWS, OCI, and other cloud providers.
- Driving system observability, fault isolation, and resilience engineering, ensuring services meet strict availability and latency SLAs.
- Providing technical mentorship to Dev Ops engineers and influence best practices in automation, monitoring, and release engineering. Champion a culture of data-driven reliability through postmortems, SLIs/ SLO's, and continuous performance optimization.
- 15+ years in Dev Ops, SRE for large-scale, production systems. successful hands-on background in Linux systems, networking, and distributed systems.
- Possess experience operating and design low-latency, high-throughput backend services at global scale. Knowledge of media or real-time communication systems (e.g., MMR, WebRTC).
- Recognize knowledge of TCP/IP, routing, DNS, load balancing, and packet capture tools. Familiarity with colocation data center operations, including hardware provisioning and troubleshooting.
- Demonstrate experience with Terraform, Ansible, Kubernetes, Docker, and modern CI/CD pipelines. successful problem-solving, debugging, and systems-level design skills
- Occasional weekend work may be required
- Ability to work across the globe or multiple time zones
Minimum: $
Maximum: $
In addition to the base salary and/or OTE listed Zoom has a Total Direct Compensation philosophy that takes into consideration; base salary, bonus and equity value.
Note:
Starting pay will be based on a number of factors and commensurate with qualifications & experience.
We also have a location based compensation structure; there may be a different range for candidates in this and other locations
Anticipated Position Close Date:
09/18/26
Ways of WorkingOur structured hybrid approach is centered around our offices and remote work environments. The work style of each role, Hybrid, Remote, or In-Person is indicated in the job description/posting.
BenefitsAs part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways. Learn for more information.
About UsZoomies help people stay connected so they can get more done together. We set out to build the best collaboration platform for the enterprise, and today help people communicate better with products like Zoom…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).