Director Cloud Operations & Support
The total target compensation (TTC) range, including salary and target bonus, is $159.300 - $247,800 . As the maximum TTC range exceeds $200,000, it is exempt from Ontario’s requirement to post a compensation range no greater than $50,000. This TTC range is applicable to permanent roles only. The actual base salary offered within this range will be determined by the successful candidate’s skills and experience, as well as internal equity.
Empire Life is looking to hire a The Director Cloud Operations & Support to join our IT team! We are actively seeking candidates to fill a current, open position
The Director Cloud Operations & Support is responsible for improving operational resiliency across the enterprise while leading the team through major technological shifts, including AI integration and the shift to cloud-first operations. The strategy is cornerstoned by the mandate to provide excellent internal service and reliability. The Director will establish a holistic system monitoring across all technologies, driving a proactive cultural shift, and embedding modern Site Reliability Engineering (SRE) principles.
Success requires a pragmatic balance between SRE automation and established processes like ITIL to ensure both high-velocity reliability and operational rigor. The Director will report to the Vice-President of IT Back Office.
Why pursue this opportunity
Our mission - Provide expertise and intelligent solutions to help Canadians navigate life with confidence.
Join a transforming business - we are a medium-sized Canadian company in a stable industry which is rapidly transforming.
Impactful work - get the opportunity to work on meaningful projects that have a positive impact on our customers, our company, and society as a whole.
Hone your skills - this is an opportunity that allows for you to grow your technical, and functional skills.
What you’ll be working on
Leadership & Reliability Engineering
Define Strategic Direction:
Establish the overarching vision, long-term strategic roadmap, monitoring, and operational excellence across the entire technology ecosystem.
Implement Key Frameworks:
Champion the adoption of SRE practices by creating and managing frameworks for Service Level Objectives (SLOs) and Error Budgets to drive accountability and continuous improvement in service reliability.
Holistic System Coverage:
Ensure the monitoring operating model provides comprehensive, real-time insight and coverage across all systems, from front-end applications to back-end infrastructure.
Stakeholder Management & Audit Interface:
Serve as the primary interface with internal Audit and Compliance teams to ensure the health, stability, and control of all operational systems and processes are clearly understood and accurately documented.
Operational Transformation & Transition
Batch Process Modernization:
Guide the crucial optimization of legacy overnight batch processes as they are migrated and modernized into efficient, cloud-native architectures.
Incident Management:
Evolve and govern the standardized, high-velocity incident management process, focusing on rapid response, clear communication, and conducting blameless post-mortems to ensure system resilience and prevent recurrence.
Toil Reduction & Automation:
Systematically identify, measure, and prioritize toil across engineering and operations teams, driving automation efforts to eliminate manual, repetitive work and free up engineering capacity for innovation and strategic projects.
Disaster Recovery (DR) Alignment:
Actively assist with and inform the Disaster Recovery (DR) strategy based on reliability data and performance trends derived from the observability framework.
Shift to Proactive Operations:
Lead the transformation of the helpdesk function to a proactive, agile operations team, focusing on incident prevention, rapid resolution, and continuous learning.
Team Development & Mentorship:
Drive a necessary culture shift from reactive, batch-centric management to a proactive, cloud-first, full-stack observability mindset. This includes mentoring managers and technical teams on SRE practices and agile operations.
Continuity & Knowledge Transfer:
Develop robust continuity plans and documentation processes to ensure resilience and consistency of operational knowledge across team members, mitigating single points of failure.
Governance & Tooling:
Define and govern the implementation of new tools, processes, and automation necessary to support scalable, real-time performance and reliability insight.
10+ years in IT, with 5+ in a senior operational leadership role. Proven track record in establishing SLOs (Service Level Objectives), driving toil reduction, and running high-availability cloud environments. Focus is on leveraging engineering practices to achieve operational excellence.
Cloud operational/security certs (e.g., AWS Certified Sys Ops Administrator or Google Cloud Professional Cloud Dev Ops Engineer).
Certifications in SRE/Dev Ops…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: