Site Reliability Engineering Lead
Listed on 2026-08-15
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Site Reliability Engineering Lead
The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams. Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime. The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks.
Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.
Essential duties and responsibilities include implementing software architecture and engineering approaches for complex initiatives, adopting and refining advanced software engineering standards, collaborating with senior engineers and product partners, leading the technical design and implementation of scalable, secure, and highly available software solutions, troubleshooting and resolving complex technical issues, providing technical guidance and training to other engineers, evaluating emerging technologies, contributing to the development of long-term technical goals, and leading large or complex initiatives.
Required qualifications include a bachelor's degree in computer science, software engineering, or related field, minimum of 7 years of professional experience in software development, deep knowledge of multiple programming languages, software architecture, and design principles, and deep understanding of software development lifecycle, testing, deployment, and security practices.
Preferred qualifications include an advanced degree in computer science or related technical discipline, professional certifications such as Certified Software Development Professional (CSDP) or equivalent, deep expertise in cloud-native architectures, microservices, container orchestration, and Dev Ops, strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management, 7+ years of experience in Site Reliability Engineering, Dev Ops, Platform Engineering, or Infrastructure Operations, deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling, proficiency with automation and scripting languages (Python, Go, Power Shell, Ansible), strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring, proven leadership in major incident management and cross-team technical coordination, strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns, excellent communication skills, including executive-level situational awareness during critical incidents, and demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).