Site Reliability Engineer
Job in
Durham, Durham County, North Carolina, 27709, USA
Listed on 2026-09-03
Listing for:
Cisco
Full Time
position Listed on 2026-09-03
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Site Reliability Engineer (SRE)
As a Site Reliability Engineer (SRE) supporting backend services for Cisco's SaaS collaboration products, you will play a critical role in delivering reliable, scalable, and resilient experiences across calling, messaging, meetings, and contact center solutions. Your work will directly impact the availability, performance, and quality of services used by millions of users globally.
Specific responsibilities include:
- Own the deployment and operation of critical collaboration services across cloud and hybrid environments, driving reliability and scalability.
- Design, evolve, and optimize CI/CD pipelines and automation, including AI-first tooling for deployment, monitoring, and incident response.
- Lead incident response for complex production issues, perform root cause analysis, and drive systemic reliability and performance improvements.
- Use observability data to guide capacity planning, scaling strategies, and resource optimization across services.
- Define and champion operational best practices, documentation standards, and a culture of reliability and operational excellence.
Minimum qualifications:
- Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience) with 5+ years in Site Reliability Engineering, Cloud Operations, or Systems Engineering.
- Strong hands-on experience operating production services using Docker and Kubernetes in cloud or hybrid environments.
- Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash) to build automation and operational tooling.
- Experience with monitoring, observability, and incident response in production environments, including on-call participation and post-incident reviews.
- Working knowledge of Linux systems, networking, distributed systems, CI/CD pipelines, infrastructure-as-code, and Git-based workflows.
Preferred qualifications:
- Experience operating large-scale, globally distributed SaaS platforms.
- Familiarity with hybrid cloud environments and multi-region deployments.
- Experience applying AI-assisted or automation-first approaches to SRE tooling and workflows.
- Strong written communication skills for creating clear operational documentation and runbooks.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×