×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Minneapolis, Hennepin County, Minnesota, 55400, USA
Listing for: RBC Capital Markets, LLC
Full Time position
Listed on 2026-07-18
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 80000 - 140000 USD Yearly USD 80000.00 140000.00 YEAR
Job Description & How to Apply Below

Job Description

RBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and platforms that support the wealth management business. Working at the intersection of software engineering, cloud-native operations, observability, and automation, the team plays a central role in delivering reliable digital services for both internal users and clients.

As a Senior Site Reliability Engineer, you will bring an engineering-first mindset, strong operational judgment, and a passion for automation to improve system reliability  will work closely with development, infrastructure, platform, and support teams to build modern observability practices, improve incident response, strengthen reliability engineering standards, and drive the evolution toward intelligent, self-healing operations.

This role is ideal for a hands‑on engineer who is equally comfortable improving production resilience, building automation, defining service‑level objectives, and shaping the future of AI‑enhanced operations. You will help design and implement scalable SRE solutions across the technology estate using tools and platforms such as Elasticsearch, Ansible, Git Hub Actions, Dynatrace, Pager Duty, Moogsoft, Kubernetes, Open Shift, Kafka, and emerging AIOps capabilities.

What

will you do?
  • Build and enhance the SRE product base – develop intelligent monitoring, alerting, reliability testing, anomaly detection, and automated remediation capabilities.
  • Implement modern observability practices – deploy metrics, logs, traces, dashboards, and actionable alerting across supported applications.
  • Design ML‑based anomaly detection and self‑healing solutions – shift from reactive to predictive operations with automated issue remediation and appropriate governance controls.
  • Standardize telemetry and instrumentation – improve visibility, coverage, and correlation of operational signals across platforms.
  • Automate operational workflows – use Ansible, Git Hub Actions, and scripting (Bash, Python, Power Shell) to streamline platform tasks and develop custom tooling.
  • Define and track service health metrics – establish and improve SLIs, SLOs, error budgets, and evolve runbooks into automation‑first remediation patterns.
  • Partner with development teams – ensure applications meet reliability and performance standards before and after deployment through close collaboration.
  • Lead incident and problem management – troubleshoot production issues across all layers, participate in on‑call rotation, and drive root‑cause analysis and corrective actions.
  • Drive continuous improvement – identify opportunities to simplify, automate, and modernize operations using engineering and AI‑driven approaches.
What do you need to succeed? Must‑have
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, Dev Ops, Platform Engineering, or Systems Engineering roles with strong operational depth.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Strong experience with infrastructure automation and configuration management, particularly Ansible.
  • Strong scripting and automation skills in Bash, Python, Power Shell, or similar languages.
  • Hands‑on experience with modern reliability and observability tooling such as Elasticsearch, Dynatrace, Git Hub, Kubernetes, Open Shift, Kafka, Pager Duty, Moogsoft, or related platforms.
  • Strong understanding of production operations, incident management, root‑cause analysis, and reliability engineering practices.
  • Experience defining and operating SLIs, SLOs, alerting strategies, and service health metrics.
  • Knowledge of cloud‑native and distributed systems concepts, including resiliency, scalability, fault isolation, and performance tuning.
  • Understanding of AIOps, AI/ML concepts, or intelligent automation as applied to observability and operations.
  • Ability to work across teams, influence engineering practices, and communicate clearly with technical and non-technical stakeholders.
Nice‑to‑have
  • Experienc…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary