Senior Manager, Site Reliability & Infrastructure Engineering
Listed on 2026-08-13
-
IT/Tech
Disaster Recovery IT, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Business Continuity
Individually we are people, but together we are Aviva. Individually these are just words, but together they are our Values – Care, Commitment, Community, and Confidence.
At Aviva Canada
, we put people first, our employees, our customers, and our communities. We’re proud of a culture built on care, inclusion, and collaboration, where your voice matters and your growth is supported. We’re not just about insurance; we’re about making a real difference by protecting what matters most.
The Senior Manager, Site Reliability & Infrastructure Engineering will lead Aviva Canada’s evolution toward an engineering-led reliability and resilience capability across critical applications, platforms and services. This role will work across on-premises, AWS, SaaS/vendor-hosted and hybrid application environments, partnering with Engineering Operations, Cloud & Platform Engineering, Application Engineering, Cybersecurity, Operational Resilience, Business Continuity, Enterprise Architecture, Vendor Management and key partners including AWS, DXC, CGI, Snowflake and other SaaS providers.
The focus is to shift Aviva from periodic recovery and resilience testing to proactive, measurable reliability engineering. This uses observability platforms like Dynatrace, incident response tools such as Pager Duty or equivalent, and modern recovery platforms like Rubrik to improve service availability, operational insight, recovery readiness, and customer outcomes.
- Establish and mature SRE practices across critical applications and platforms, including SLIs, SLOs, SLAs, service health indicators, post-incident reviews and continuous reliability improvement.
- Partner with application, platform, infrastructure and vendor teams to embed reliability requirements into design, development, operational acceptance, release and production support processes.
- Drive improvements in incident response, problem management, root cause analysis, alert quality, critical issue workflows, service health reporting and reduction of repeat incidents.
- Improve Infrastructure Ops via automation, self-healing patterns, runbook improvements, AI-assisted operations and repeatable engineering practices.
- Use observability insights to improve infrastructure availability, performance, capacity planning, & operational readiness.
- Maintain and evolve Aviva Canada’s technology resilience capability across on-premises platforms, AWS, critical applications, integration services and SaaS/vendor-hosted platforms.
- Ensure resilience planning validates end-to-end recoverability, including application, data, integration, identity, network, platform, cloud and vendor dependencies.
- Support annual BCP, DR and technology resilience testing, including RTO/RPO validation, dependency mapping, recovery sequencing, test evidence, gap management and remediation tracking.
- Own and mature backup and cyber recovery practices, including hands‑on use of Rubrik or a comparable enterprise data protection platform for backup policy management, immutability, restore validation, coordinating recovery procedures, reporting, evidence capture and operational support for critical workloads.
- Partner with Cybersecurity on ransomware and destructive cyber event readiness, including clean recovery, backup integrity validation, isolated recovery environments, cyber recovery vault operations and restoration of critical services.
- Ensure resilience and recovery requirements are embedded into architecture, cloud migration, vendor onboarding, operational readiness, change delivery and service governance.
- Lead, manage and coach group of engineers working across reliability, observability, production support, technology resilience and recovery practices.
- Build a clear operating model for SRE and technology resilience ownership across application teams, infrastructure teams, cloud/platform teams, cybersecurity, operational resilience and third‑party providers.
- Define reliability and resilience reporting for senior leaders, including service health, SLO performance, incident trends, MTTR, alert quality, recovery readiness, open…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: