Senior DevOps & SaaS Operations Engineer
Listed on 2026-08-27
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Cybersecurity
Senior Dev Ops & SaaS Operations Engineer
Xage Security
August 21, 2026
Location:
Palo Alto, CA
Cyberattacks on critical infrastructure, government, and private enterprises are increasing in scale, speed, and impact. Xage Security is the leader in Zero Trust for high-stakes environments, helping organizations protect what matters most. With Xage, organizations can stop cyber threats, strengthen resilience, simplify security, and move faster with confidence.
We have built tremendous momentum across governments and commercial enterprises around the world, and it’s just the beginning. Recognized by Forbes as one of America’s Best Startup Employers, Xage prioritizes creativity, collaboration, and innovation in pursuit of our mission. We are headquartered in Palo Alto, CA and have global teams across North America and EMEA.
We’re passionate about solving problems that have positive, real-world consequences for the lives of everyday people. We hope you’ll join us in protecting what matters most and helping organizations operate with confidence in an increasingly complex threat landscape.
About the RoleWe are seeking a Senior Dev Ops & SaaS Operations Engineer to design, automate, and scale the cloud infrastructure and operational pipelines powering our hosted Zero Trust AI Gateway platform. As our SaaS solution expands across enterprise customers, ensuring high availability, seamless customer deployments, automated upgrades, and proactive monitoring across both cloud and edge environments is critical.
Today, our platform runs on a VM-based and containerized architecture (Docker, systemd, cloud compute instances). In this role, you will own the current infrastructure and automating deployments, tenant management, and monitoring, while designing the architectural bridge and migration path toward Kubernetes (AKS/EKS/GKE) as our scale demands it.
Key Responsibilities- Customer Deployment & Onboarding Automation: Build and maintain Infrastructure-as-Code (Terraform, Packer) and automated provisioning workflows to seamlessly deploy, isolate, and configure customer environments across cloud VMs and container runtimes.
- Zero-Downtime Upgrade Orchestration: Architect safe, automated upgrade and rollback pipelines for both hosted cloud control planes and customer-side edge gateways/agents, ensuring continuous operations across version shifts.
- Kubernetes Migration Strategy: Lead the future-state container orchestration roadmap—designing, prototyping, and executing the transition from VM/Docker deployments to Kubernetes (EKS/GKE) without interrupting customer SLAs.
- Monitoring, Observability & Alerting: Design and manage unified observability platforms (e.g., Prometheus, Grafana, Datadog, Click House, Open Telemetry) to track system health, gateway latency, error budgets, and service SLIs/SLOs.
- Ongoing Maintenance & Site Reliability: Drive operational excellence, capacity planning, backup/disaster recovery, patch management, and incident response automation to maintain enterprise-grade uptime SLAs (99.99%).
- Cloud & Edge Security Infrastructure: Implement security best practices across VM images, container registries, secret management systems, and network perimeter controls across AWS/GCP/Azure environments.
- Engineering Degree or equivalent experience
- Production SaaS
Experience:
5+ years of experience in Dev Ops, Site Reliability Engineering (SRE), or Cloud Operations managing commercial SaaS platforms. - Infrastructure as Code (IaC) & Cloud: Expert-level mastery of Terraform (or Open Tofu), Packer, and deep experience operating within primary cloud providers (
AWS, GCP, or Azure ). - VM & Container Management: Hands-on experience deploying, managing, and hardening VM-based workloads (Azure VM/EC2/GCE) along with containerization (
Docker, Docker Compose ). - Kubernetes Expertise: Hands-on experience with Kubernetes (EKS/GKE) and Helm, with a clear understanding of how to architect containerized applications for future Kubernetes migration.
- CI/CD & Release Engineering: Proven track record building robust deployment pipelines (Git Hub Actions, Ansible, Packer) featuring blue-green deployments, canary…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).