×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Mississauga, Ontario, Canada
Listing for: 0000050007 Royal Bank of Canada
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Job Description & How to Apply Below

Job Description

What is the opportunity?RBC Insurance Technology is seeking to hire a Senior Site Reliability Engineer for its Insurance Technology Platform Support team. The Insurance Technology Platform Support Team is a specialized unit dedicated to ensuring the optimal performance, availability, and resilience of IT applications used in the insurance line of business. With a unique blend of technical expertise and industry-specific knowledge, this team plays a critical role in ensuring the seamless operations of digital services that cater to both the business's internal and external stakeholders.

As a Senior Site Reliability Engineer, you will bring the engineering mindset of bold ambition, curiosity and outcome focus to ensuring the performance and reliability of our systems. This role calls for a dynamic individual who excels in a collaborative environment, working with cross-functional teams to implement best practices for observability, monitoring, logging, alerting, and automation. As we evolve toward AI-driven autonomous operations, you will play a key role in transitioning from traditional reactive incident response to intelligent, self-healing systems.

This role will be responsible for the development, implementation, and support of Site Reliability Engineering (SRE) solutions for applications supported by RBC Insurance Technology. You'll leverage your proficiency in Elasticsearch, Ansible, Git Hub Actions, Moogsoft, Pager Duty, Dynatrace, and emerging AIOps platforms to build and maintain robust automation, intelligent observability, and AI-enhanced SRE tooling.
What will you do?
  • Contribute to the SRE product base (intelligent monitoring, alerting, machine learning anomaly detection, Agentic AI self-healing, reliability testing)

  • Implement and enhance AI-driven monitoring and intelligent observability capabilities across supported applications

  • Design and implement ML-based anomaly detection pilots, transitioning from rule-based to predictive alerting

  • Architect and develop Agentic AI self-healing solutions that autonomously remediate common incidents

  • Design human-AI workflows that balance automation efficiency with appropriate human oversight and governance

  • Standardize application telemetry data to increase coverage of signal types, building the foundation for advanced AI/ML capabilities

  • Contribute to centralization of observability and monitoring backends for advanced telemetry correlation

  • Collaborate with cross-functional teams to implement best practices for monitoring, logging, and incident response, driving a proactive stance on system health

  • Implement and manage automation processes with Ansible and Git Hub Actions to streamline operational tasks

  • Develop and maintain custom tooling and automation scripts in languages like Bash, Python, and Power Shell to enhance operational efficiency and system reliability

  • Work closely with development teams to understand code changes and their impact on the production environment, ensuring that new releases meet our reliability standards

  • Actively contribute to the definition and tracking of SLIs, SLOs, and other critical metrics, refining our alerting and monitoring strategies accordingly

  • Evolve runbooks into automated remediation workflows and Agentic AI automation, reducing manual intervention

  • Create and refine custom tooling and automation scripts using languages such as Bash, Python, and Power Shell, supporting the infrastructure's scalability and reliability needs

  • Support deployments by advocating for reliability and performance improvements based on industry trends and company objectives

  • Participate in incident management and problem management for applications in scope and contribute to RCA Action items fulfillment

  • Validate and govern AI outputs to ensure compliance with financial services regulations and maintain human accountability for AI-driven decisions

  • Drive transformation by continuously looking for ways to automate existing processes and adopt intelligent operations

  • Debug production issues across services and levels of the stack and provide primary operational support

  • Perform production support role, including off-hours support (as part of an on-call rotation)

  • Must-have
  • 4+ years of SRE or Systems Engineering experience with strong technical expertise

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience

  • Expertise in infrastructure-as-code and configuration management, particularly Ansible

  • Advanced scripting capabilities in Bash, Python, Power Shell, or other similar languages

  • In-depth knowledge of tools such as Elasticsearch, Ansible, Git Hub, Open Shift, Kubernetes, Dynatrace, Kafka, and their role in system reliability

  • Knowledge of creating, maintaining, and alerting on SLIs, SLOs, and other reliability metrics

  • Understanding of AI/ML concepts and their application to observability and operations (AIOps)

  • Experience with or strong interest in intelligent monitoring, anomaly detection, and automation technologies

  • Ability to…

  • Position Requirements
    10+ Years work experience
    Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
    To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary