Sr Engineer, Restaurant Systems - Edge Platform & Kubernetes Operations
Listed on 2026-09-15
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Key Responsibilities Job Description Senior Edge Platform Engineer Location
Dallas, TX (Hybrid) or Full Remote Travel:
Up to 20% to restaurant, lab, and vendor locations
The Senior Engineer, Edge Platform & Kubernetes Operations is responsible for the design, implementation, support, and continuous evolution of the company's highly available edge computing platform that will be servicing more than 1,200 restaurant locations. This role will own the operational health, scalability, security, and reliability of a fleet of K3s-based edge clusters that host critical restaurant technology workloads, including POS, payment processing, kitchen systems, third-party integrations, and other business-critical applications.
This engineer will work at the intersection of cloud infrastructure, Kubernetes, virtualization, storage, networking, observability, and automation to ensure restaurant operations remain resilient even during hardware failures, connectivity disruptions, and software upgrades. The role partners closely with Platform Engineering, Restaurant Systems, Infrastructure, Cybersecurity, Service Desk, and external vendors to deliver a fault-tolerant, enterprise-grade edge ecosystem.
- Design, implement, and support a fleet of highly available K3s clusters deployed across 1,200+ restaurant locations.
- Develop and maintain edge computing architectures that support local autonomy while integrating with centralized Google Cloud Platform (GCP) and/or other cloud service(s) platforms.
- Engineer solutions for automatic failover, disaster recovery, workload mobility, quorum management, and node resiliency.
- Define hardware, storage, and networking standards for edge deployments.
- Serve as the technical leader for Kubernetes operations, upgrades, lifecycle management, and cluster health.
- Manage containerized and virtualized workloads running on Kubernetes and Kube Virt.
- Establish operational standards for cluster provisioning, scaling, patching, and decommissioning.
- Improve platform reliability through infrastructure automation and self-healing capabilities.
- Design and support shared storage architectures using technologies such as Rook/Ceph.
- Monitor cluster quorum, data replication, storage performance, and recovery procedures.
- Lead testing and validation of failover scenarios, node outages, and disaster recovery processes.
- Ensure stateful workloads maintain availability and data integrity during failure events.
- Integrate edge computing platforms with Google Cloud services including:
- GKE
- Cloud Monitoring
- Cloud Logging
- Artifact Registry
- Compute Engine
- Cloud KMS
- Design secure management and observability paths between restaurant edge clusters and cloud-hosted services.
- Optimize cloud consumption and operational efficiency.
- Build and maintain enterprise observability solutions leveraging Prometheus, Grafana, monitoring platforms, and centralized telemetry pipelines.
- Establish service-level indicators (SLIs), service-level objectives (SLOs), and reliability metrics.
- Drive root cause analysis and post-incident reviews for platform outages and service degradations.
- Partner with support organizations to develop operational runbooks and escalation procedures.
- Build Infrastructure-as-Code and Git Ops-based deployment models.
- Automate edge cluster deployment, onboarding, upgrades, and compliance validation.
- Develop CI/CD processes supporting edge and cloud-native applications.
- Reduce operational overhead through orchestration, scripting, and platform automation.
- Implement secure-by-design platform controls across compute, storage, networking, and cloud integrations.
- Work closely with Security teams on CIS hardening, vulnerability remediation, certificate management, and platform compliance.
- Ensure restaurant systems meet company security and data protection standards.
- Act as a subject matter expert for Kubernetes, edge computing, and distributed systems.
- Mentor engineers and support teams on cloud-native technologies and operational best practices.
- Evaluate emerging technologies and develop future-state architecture recommendations.
- Participate in vendor evaluations, proof-of-concepts, and platform roadmap planning.
Required Qualifications
- 8+ years of infrastructure, cloud, systems engineering, or platform engineering experience.
- 5+ years managing…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).