Lead, reliability engineering
Date limite pour présenter sa candidature :
09/03/2026
Adresse :
33 Dundas Street West
Groupe de famille d'emploi :
Technologie
Transformation Mandate
This role exists to fundamentally shift Commercial Banking technology operations from reactive support to an AI-enabled, engineering-driven reliability model at scale.
The Lead will own the outcome of this transformation, challenging the status quo and driving a step-change in automation, observability, and operational excellence.
Role OverviewWe are looking for a high-performing, results-driven senior technology leader to take on the role of Lead, Technology Support & Reliability Engineering. This is a critical leadership position responsible for driving the reliability, resiliency, and modernization of technology platforms supporting Commercial Banking.
The successful candidate will set the vision and lead the transformation of traditional support into a modern, engineering-led, highly automated operating model. This includes the aggressive adoption of Site Reliability Engineering (SRE), AIOps, and a forward-looking No
Ops philosophy to reduce toil, increase automation, and enable self-healing, resilient systems at scale.
This role demands a proven high performer who consistently owns the outcome, questions the status quo, and delivers measurable improvements in stability, efficiency, and engineering velocity.
Key Responsibilities Reliability & Production Support Leadership- Lead L2 and L3 technology support and SRE teams across Commercial Banking platforms
- Own major incident management, escalation, and recovery processes
- Drive service stability, availability, and resiliency outcomes through the entire tech stack, including supporting functions such as infrastructure, helpdesk, data lake, etc.
Ops Transformation
- Lead the transformation toward a modern SRE-driven operating model incorporating AIOps and No
Ops principles - Drive implementation of observability frameworks, predictive analytics, and event correlation
- Enable proactive operations through intelligent automation and predictive capabilities
- Reduce operational toil through automation, self-healing systems, and intelligent workflows
- Shift teams from reactive support to engineering-led operations
- Build automated workflows and deployment pipelines
- Enable end-to-end service automation and platform abstractions
- Support cloud adoption and improve proactive operational capabilities
- Define and execute technology support strategy aligned to enterprise and Commercial Banking priorities
- Own and evolve the L2/L3 operating model
- Establish clear RACI and ownership models across engineering and operations
- Drive governance, controls, and continuous improvement
- Act as a senior technology executive and trusted advisor
- Manage cross-functional dependencies across engineering and platform teams
- Lead large-scale strategic initiatives and transformation programs
- Build and lead high-performing teams with a strong culture of accountability and ownership
- Attract, retain, and develop top engineering and SRE talent
- Drive a performance-oriented culture aligned to enterprise objectives
- Mentor and develop leaders capable of operating at scale
- Reduce Major Incident volume (P1/P2) by 30-50%
- Improve Mean Time to Restore (MTTR) by 25-40%
- Deliver sustained improvements in service availability and resiliency
- Eliminate 95% of manual operational toil
- Increase automated incident resolution through AIOps
- Reduce human intervention in steady-state operations
- Establish end-to-end observability coverage
- Implement event correlation, predictive alerting, automated triage
- Improve alert quality by reducing noise
- Transition to SRE-led model
- Define L2/L3 ownership
- Shift to proactive incident prevention
- Build high-performance engineering culture
- Increase productivity and engagement
- Develop senior SRE capability
- Both cloud (AWS/Azure) and on prem deployment and support
- Deep observability experience is a must, preferably with some exposure to Dynatrace
- Expertise in incident, problem, change management
- Demonstrated success in automation (Ansible a plus) and AIOps
- Experience with Agile methodologies
- Financial services experience
- SRE/Dev Ops/platform engineering knowledge
- ITIL Foundations or better
- High-performance, results-oriented mindset
- Owns the outcome and challenges status quo
- Strategic thinker
- Strong communicator
- Passion for innovation and automation
This role is central to transforming technology support into a modern, AI-enabled, highly automated reliability organization, impacting customer experience, operational risk, and delivery speed.
Salaire :$ - $
Type de…To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: