Lead Principal Site Reliability Engineer
Listed on 2026-08-01
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description
Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks.
Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.
Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks.
Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.
Key Responsibilities Reliability Strategy and Technical Leadership
- Define and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms.
- Establish reliability standards, engineering practices, and operational readiness requirements across multiple teams.
- Serve as a senior technical authority for system reliability, scalability, resilience, performance, and production operations.
- Influence architecture and design decisions to ensure systems are supportable, observable, fault tolerant, and capable of meeting availability objectives.
- Identify systemic reliability risks and lead cross-functional initiatives to address them.
- Provide technical direction and mentorship to site reliability engineers, software engineers, platform engineers, and operations teams.
- Lead technical reviews and promote consistent engineering practices across the organization.
- Define and implement service-level indicators, service-level objectives, error budgets, and operational health metrics.
- Develop comprehensive monitoring, logging, tracing, alerting, and observability strategies.
- Improve the quality and actionability of alerts while reducing unnecessary operational noise.
- Establish dashboards and reporting mechanisms that clearly communicate service health, performance, capacity, and risk.
- Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives.
- Design and implement automation that reduces manual intervention, operational toil, and human error.
- Build or enhance tools for deployment, configuration management, infrastructure provisioning, incident response, and service recovery.
- Promote infrastructure-as-code, policy-as-code, automated testing, and repeatable deployment practices.
- Partner with development teams to improve continuous integration and continuous delivery pipelines.
- Develop self-healing and automated remediation capabilities where appropriate.
- Contribute production-quality software and reusable platform components using modern programming and scripting languages.
- Provide technical leadership during complex, high-severity production incidents.
- Coordinate diagnosis, containment, recovery, and stakeholder communication during service disruptions.
- Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms.
- Identify recurring failure patterns and develop long-term engineering solutions.
- Improve incident-management processes, escalation procedures, runbooks, and recovery playbooks.
- Participate in an on-call rotation or provide senior escalation support for critical services, as required.
- Lead capacity planning, performance analysis, load testing and demand forecasting for critical platforms.
- Identify performance…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).