Manager, Site Reliability Engineering (SRE
Listed on 2026-08-24
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Project Manager, Systems Engineer
Position: Manager, Site Reliability Engineering (SRE)
Location: Toronto
Salary: $155,000 - $165,000
Posting Type: Open vacancy
Our client is looking for an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and resiliency of business-critical applications and platforms. This is a newly created position reporting to the Director, SRE & Dev Ops. The successful candidate will combine people leadership with strong hands-on technical expertise, helping mature an evolving SRE function while supporting a large portfolio of critical applications.
Our client is looking for someone who can lead and develop an SRE team while remaining technically credible and hands-on across cloud, automation, observability, infrastructure and production reliability.
- Lead, coach and develop a team of Site Reliability Engineers.
- Own the reliability, availability, scalability and performance of assigned business-critical applications and platforms.
- Help mature SRE practices including SLIs, SLOs, error budgets, incident response, post-incident reviews and continuous reliability improvement.
- Oversee production operations, incident management, escalation handling and problem management.
- Drive improvements in observability, monitoring and operational visibility.
- Partner with Development, Dev Ops, Infrastructure, Security and Incident Management teams.
- Reduce operational toil through automation and self-healing capabilities.
- Lead capacity planning, resiliency testing and disaster recovery readiness.
- Act as a senior technical escalation point during major production incidents.
- Establish effective runbooks, documentation standards and knowledge-management practices.
- Help develop reliability roadmaps for the applications within your portfolio.
- Support hiring and workforce planning as the SRE organization continues to evolve.
- 8+ years of experience in senior technical roles supporting complex or distributed systems.
- 3+ years of people leadership experience.
- Strong hands-on knowledge of at least one major cloud platform, with Azure preferred; AWS is also acceptable.
- Strong knowledge of Kubernetes.
- Near-expert capability with Infrastructure as Code, automation and observability.
- Strong experience with enterprise observability tools;
Dynatrace is preferred, with Datadog, New Relic, App Dynamics or similar also relevant. - Linux knowledge – this is an important technical requirement.
- Strong scripting and/or programming capabilities.
- Strong understanding of SRE principles including SLOs, SLIs, incident management, resiliency and toil reduction.
- Ability to operate as both a people leader and hands-on technical leader.
- Excellent communication and stakeholder-management skills.
- Experience with in payments, fintech or other highly regulated environments.
- Exposure to PCI DSS and/or NIST frameworks.
- Experience working with change-management and compliance processes.
- Familiarity with SDLC best practices.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: