Engineer lead, Cloud Foundation Services Platform; Kubernetes- ST; Seattle WA
Listed on 2026-09-03
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Cloud Engineer – Lead
At Starbucks, our mission is to inspire and nurture the human spirit – one person, one cup, and one neighborhood at a time. Starbucks Technologists work to achieve this mission through the use of cutting-edge technology delivered to our partners, customers, stores, roasters, and global communities.
This job contributes to Starbucks success by delivering high-quality, reliable, and stable technologies and security capabilities in support of the Starbucks Engineering Platform. This position is accountable for the installation, configuration, monitoring, analysis, maintenance, and technical support of the platform.
We are seeking a Cloud Engineer – Lead with deep expertise in Azure Kubernetes Service (AKS) and cloud-native platform engineering to design, build, and operate a secure, scalable, and cost-efficient Kubernetes-based platform on Microsoft Azure. This role will serve as a hands-on technical leader responsible for enabling development teams through reliable platforms, automation, and modern Dev Ops and Git Ops practices.
The ideal candidate is an AKS and Kubernetes expert with strong working knowledge of Azure PaaS services, service mesh technologies, and continuous delivery systems, and who brings a strong focus on security, reliability, and cost optimization.
Key ResponsibilitiesTechnical Leadership & Collaboration
- Communicate complex platform and Kubernetes architecture decisions clearly to both technical and non-technical stakeholders
- Establish strong cross-functional partnerships with application, security, networking, and business teams
- Act as a technical leader and mentor for Kubernetes, AKS, and Azure platform engineering best practices
- Partner with technology vendors and open-source communities to deliver against business and platform objectives
AKS & Platform Engineering
- Design, implement, and operate Azure Kubernetes Service (AKS) platforms at scale
- Build and maintain a secure, multi-tenant Kubernetes platform with a focus on:
- Reliability, performance, and scalability
- Environment consistency across dev, test, and production
- Leverage Azure-native services and PaaS offerings to support application and platform needs
- Own Kubernetes cluster lifecycle management, upgrades, capacity planning, and operational health
Kubernetes Ecosystem & Service Mesh
- Design and operate Kubernetes-native solutions using:
- Argo CD / Argo Workflows for Git Ops and continuous delivery
- Istio (or similar service mesh) for traffic management, security, and observability
- Implement best practices around ingress, egress, service discovery, and policy enforcement
- Enable secure communication, workload identity, and secrets management across the platform
Automation, Infrastructure as Code & Git Ops
- Drive an automation-first approach using:
- Terraform, Azure Bicep, and Crossplane for infrastructure and platform provisioning
- Python for scripting, tooling, and automation
- Enable CI/CD and Git Ops workflows using Azure Dev Ops, Git-based pipelines, and modern delivery patterns
- Engineer standardized, repeatable build and release processes for platform and application teams
Security, Compliance & Cost Optimization
- Design platforms that are secure by default, following Zero Trust and least-privilege principles
- Ensure all implementations align with Information Security policies and compliance requirements (e.g., PCI)
- Continuously evaluate and optimize platform cost using cloud-native tooling and usage analysis
- Implement system configurations and baselines that support secure software development and operational best practices
Observability & Operations
- Implement deep telemetry, logging, and monitoring across Kubernetes and Azure PaaS platforms
- Deploy and maintain observability solutions to support performance, reliability, and troubleshooting
- Ensure high availability and operational continuity for mission-critical services
- Support and operate 24x7 production environments, enabling automated recovery wherever possible
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent professional experience)
- 8–10 years of professional industry experience in cloud, platform, or software engineering
- 2…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).