Senior Observability Engineer
Listed on 2026-07-20
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cybersecurity, Cloud Computing: Infrastructure & Operations
Job Description
- Design and implement observability solutions using industry-leading platforms, establishing logging standards that enable comprehensive system visibility across all systems
- Create and maintain monitoring dashboards that provide actionable insights into system health and performance, partnering with platform and application teams to integrate observability into architecture
- Evaluate and recommend observability tools and vendors to ensure the organization has access to best-in-class solutions
- Analyze logs, metrics, and traces to proactively identify system issues, performance bottlenecks, and translate observability data into insights about user patterns and system behavior
- Develop predictive monitoring strategies using AI tools to detect anomalies, identify emerging trends, and prevent incidents before they occur
- Conduct "what if" analysis using AI capabilities to model potential scenarios and their impact on system performance
- Apply machine learning-based anomaly detection to identify issues proactively and use predictive analytics to forecast system behavior and prevent failures
- Collaborate with cross-functional teams (Engineering, Dev Ops, Security, production support, product) to communicate complex observability concepts to both technical and non-technical stakeholders
- Lead observability initiatives and mentor junior team members on best practices while facilitating design and problem-solving discussions across the organization
- 5+ years of experience with observability tools (ELK Stack, Dynatrace, Prometheus, Grafana, Open Telemetry, Jaeger, Aternity and Moog or similar)
- 3+ years of software engineering or infrastructure experience
- Python, Java, GO
- Query languages
- Expert-level knowledge of logging requirements and best practices for enhanced observability
- Demonstrated experience building and optimizing monitoring dashboards
- Proven ability to use observability data to proactively identify and resolve system issues
- Experience using AI tools for anomaly detection and trend analysis
- Alerting
- Linux, Unix
- Expertise with Tableau or advanced visualization tools
- Experience in financial technology and financial services environments
- Previous experience using recent AI tools such as Anthropic, OpenAI, Devin, Copilot, etc.
- Experience with Open Shift, Kubernetes and containerized applications
- Background in Dev Ops or Site Reliability Engineering (SRE)
- Experience conducting "what if" analysis and scenario modeling
- Identify networking slowness and availability issues
- A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, and stock where applicable
- Leaders who support your development through coaching and mentoring opportunities
- Ability to make a difference and lasting impact on system reliability and user experience
- Work in a dynamic, collaborative, progressive, and high-performing team
Expected salary range: $90,000-$140,000
, depending on your experience, skills, and registration status, market conditions and business needs. You have the potential to earn more through RBC’s discretionary variable compensation program which gives you an opportunity to increase your total compensation, provided the business meets its performance targets and you meet your individual goals.
- Agile SDLC
- AI Observability
- Anthropic Claude AI
- Database Queries
- Data Query Language
- Dynatrace Administration
- Elastic Stack (ELK)
- IT Monitoring
- Java (Programming Language)
- Microsoft Copilot
- Monitoring Tools
- OpenAI
- Problem Solving
- Python (Programming Language)
- Site Reliability Engineering
- Software Development
- SRE Observability
- Telemetry Monitoring
Address: 250 NICOLLET MALL:
MINNEAPOLIS
City: Minneapolis
Country: United States of America
Work hours/week: 40
Employment Type: Full time
Platform: TECHNOLOGY AND OPERATIONS
Job Type: Regular
Pay Type: Salaried
Posted Date:
Final date to receive applications:
Note:
Applications will be accepted until 11:59 PM on the day prior to the Final date to receive applications date above.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).