Senior Site Reliability Engineer
Listed on 2026-07-18
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description
RBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and platforms that support the wealth management business. Working at the intersection of software engineering, cloud-native operations, observability, and automation, the team plays a central role in delivering reliable digital services for both internal users and clients.
As a Senior Site Reliability Engineer, you will bring an engineering-first mindset, strong operational judgment, and a passion for automation to improve system reliability will work closely with development, infrastructure, platform, and support teams to build modern observability practices, improve incident response, strengthen reliability engineering standards, and drive the evolution toward intelligent, self-healing operations.
This role is ideal for a hands‑on engineer who is equally comfortable improving production resilience, building automation, defining service‑level objectives, and shaping the future of AI‑enhanced operations. You will help design and implement scalable SRE solutions across the technology estate using tools and platforms such as Elasticsearch, Ansible, Git Hub Actions, Dynatrace, Pager Duty, Moogsoft, Kubernetes, Open Shift, Kafka, and emerging AIOps capabilities.
Whatwill you do?
- Build and enhance the SRE product base – develop intelligent monitoring, alerting, reliability testing, anomaly detection, and automated remediation capabilities.
- Implement modern observability practices – deploy metrics, logs, traces, dashboards, and actionable alerting across supported applications.
- Design ML‑based anomaly detection and self‑healing solutions – shift from reactive to predictive operations with automated issue remediation and appropriate governance controls.
- Standardize telemetry and instrumentation – improve visibility, coverage, and correlation of operational signals across platforms.
- Automate operational workflows – use Ansible, Git Hub Actions, and scripting (Bash, Python, Power Shell) to streamline platform tasks and develop custom tooling.
- Define and track service health metrics – establish and improve SLIs, SLOs, error budgets, and evolve runbooks into automation‑first remediation patterns.
- Partner with development teams – ensure applications meet reliability and performance standards before and after deployment through close collaboration.
- Lead incident and problem management – troubleshoot production issues across all layers, participate in on‑call rotation, and drive root‑cause analysis and corrective actions.
- Drive continuous improvement – identify opportunities to simplify, automate, and modernize operations using engineering and AI‑driven approaches.
- 5+ years of experience in Site Reliability Engineering, Production Engineering, Dev Ops, Platform Engineering, or Systems Engineering roles with strong operational depth.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong experience with infrastructure automation and configuration management, particularly Ansible.
- Strong scripting and automation skills in Bash, Python, Power Shell, or similar languages.
- Hands‑on experience with modern reliability and observability tooling such as Elasticsearch, Dynatrace, Git Hub, Kubernetes, Open Shift, Kafka, Pager Duty, Moogsoft, or related platforms.
- Strong understanding of production operations, incident management, root‑cause analysis, and reliability engineering practices.
- Experience defining and operating SLIs, SLOs, alerting strategies, and service health metrics.
- Knowledge of cloud‑native and distributed systems concepts, including resiliency, scalability, fault isolation, and performance tuning.
- Understanding of AIOps, AI/ML concepts, or intelligent automation as applied to observability and operations.
- Ability to work across teams, influence engineering practices, and communicate clearly with technical and non-technical stakeholders.
- Experienc…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).