Technical Leader, SIte Reliability Engineer
Listed on 2026-09-12
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
This position will perform work that the U.S. government has specified can only be performed by a U.S. citizen on U.S. soil. Meet the Team
The Platform Engineering organization is responsible for building and operating the foundational cloud-native platforms that power engineering across Cisco Network Platform. Our team owns the Kubernetes ecosystem used by thousands of developers, spanning hundreds of clusters across public and on-premise environments. We provide the infrastructure, automation, observability, security, and self-service capabilities that enable product teams to build and operate services reliably are a highly collaborative team of SREs who enjoy solving complex distributed systems challenges.
Our culture emphasizes ownership, operational excellence, automation, and continuous improvement. This is an opportunity to shape the future of Kubernetes and platform engineering within Cisco Network Platform while influencing the experience of engineers across the company.
- Lead the design, implementation, and optimization of our Kubernetes platform capabilities across cloud and on-premises environments, supporting mission-critical workloads at scale.
- Drive technical initiatives that improve reliability, security, scalability, and developer productivity for hundreds of engineering teams using Infrastructure-as-Code and Git Ops.
- Partner with engineering, security, and product teams to execute platform roadmaps and deliver on company-critical reliability goals.
- Provide technical leadership during major incidents, architecture reviews, platform transformations, and strategic infrastructure investments.
- Mentor senior engineers and help raise the engineering bar through hands‑on technical leadership, code/design reviews, operational excellence, and automated remediation.
- Bachelor's degree (or equivalent experience) and 12+ years of experience designing, building, and operating large-scale distributed systems in production environments, with a proven track record of leading cross-organizational platform initiatives.
- 7+ years of experience operating Kubernetes platforms and cloud-native infrastructure in production (including multi-cluster environments) and expertise in cloud-native infrastructure across AWS, Azure, GCP, on-prem or hybrid clouds.
- Prior experience designing, operating, and troubleshooting networking architectures for large-scale Kubernetes platforms, including CNI implementations and service mesh.
- Experience with Infrastructure-as-Code and Git Ops technologies, specifically utilizing Terraform and ArgoCD.
- Experience implementing SRE practices, including SLOs/SLIs, incident management, production readiness reviews, and automated remediation.
- Strong communication and leadership skills with the ability to effectively influence engineers, architects, senior leaders, and executive stakeholders.
- Strong software engineering experience with Go, Python, ruby, or similar programming languages used for automation and platform development.
- Experience with migrating monolithic applications to a microservices architecture.
- Knowledge of foundational networking concepts (e.g., switching, routing, Wi‑Fi, IoT).
- Advanced knowledge of observability tools and implementing comprehensive monitoring strategies for distributed systems.
- Experience driving toil reduction through advanced automation and self-service developer platforms.
At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era - and beyond. We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).