Site Reliability Engineer
Listed on 2026-09-03
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
Site Reliability Engineer
Remote, US
Here at Ooma we empower people to connect in smarter ways. We do this by creating powerful communication experiences through our cloud-based platform to bring people together at work and solutions help small business owners stay connected with their customers and manage their businesses from anywhere. For larger companies we provide customized unified communications solutions to meet their unique needs. At home, we help our customers connect with their loved ones by providing the #1 rated VoIP phone service available.
We also provide them with peace of mind through our innovative smart home security solution. At Ooma, all our products and services are priced competitively, because we believe advanced technology should be accessible to all.
About the Role:
As a Site Reliability Engineer, you will leverage your extensive expertise in Linux systems, virtualization, containers, Kubernetes clusters, and CI/CD pipelines to ensure the stability and efficiency of our systems, collaborating across teams to implement best practices for infrastructure management, automated deployment, and application performance monitoring.
Deep on-premises experience is a core requirement, not a secondary consideration. Our production environment runs on our own hardware — large data centers built on hundreds of bare metal servers and VMs, with our own storage, virtualization, and physical network beneath them. You will operate comfortably at the hardware and OS layer while also owning the container and delivery platform on top of it.
What You'll Do:
- Provide expert guidance on managing large data centers, including hundreds of bare metal servers and virtual machines (VMs), ensuring optimal configuration and performance.
- Monitor and troubleshoot system performance, reliability, and availability using modern observability tools and techniques, with strong emphasis on diagnosing and resolving issues in operating systems and bare metal environments.
- Manage the server hardware lifecycle — firmware and BIOS baselines, out-of-band management, failure triage, vendor and RMA coordination, capacity forecasting, and refresh planning.
- Administer virtualization platforms and manage storage appliances across all data center locations worldwide.
- Automate bare metal and VM provisioning and OS lifecycle at scale.
- Work hands-on with data center network infrastructure — VLANs, routing, link aggregation, load balancers, and firewall rules — troubleshooting latency, throughput, and packet loss from the Linux host outward, including multi-site and colocation resiliency.
- Design, implement, and maintain scalable, reliable infrastructure using containers, Kubernetes, and microservices architecture, including clusters on bare metal and on-premises VMs.
- Oversee configuration management for consistent, reliable releases across environments, using Ansible for system configuration, patch management, and provisioning across data center infrastructure, and eliminating configuration drift across the fleet.
- Design and operate high-throughput Kafka clusters for event streaming — topics, partitions, replication, consumer lag monitoring, and disaster recovery across data center infrastructure.
- Implement name services and server management practices, including DNS, DHCP, NTP, directory and authentication services, and certificate management.
- Collaborate with development teams to influence system design choices and operational policies, and continuously evaluate and integrate new technologies, hardware platforms, and automation approaches to improve operational efficiency and reliability.
- Participate in on-call rotations supporting production systems, conduct blameless post-mortems with root cause analysis, and maintain incident response runbooks and procedures.
- Create comprehensive technical documentation — runbooks, architectural diagrams, network topology maps, rack elevations, and capacity models — and maintain knowledge bases for operational procedures and best practices.
Experience We're Looking For:
- 8+ years of experience as an SRE or in a related field, with a strong focus on production systems, containers, microservices, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).