Senior Observability Engineer
Listed on 2026-07-18
-
IT/Tech
Systems Engineer, Cybersecurity, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Title
Job Description
What will you do?
Design and implement observability solutions using industry-leading platforms, establishing logging standards that enable comprehensive system visibility across all systems.
Create and maintain monitoring dashboards that provide actionable insights into system health and performance, partnering with platform and application teams to integrate observability into architecture.
Evaluate and recommend observability tools and vendors to ensure the organization has access to best-in-class solutions.
Analyze logs, metrics, and traces to proactively identify system issues, performance bottlenecks, and translate observability data into insights about user patterns and system behavior.
Develop predictive monitoring strategies using AI tools to detect anomalies, identify emerging trends, and prevent incidents before they occur.
Conduct "what if" analysis using AI capabilities to model potential scenarios and their impact on system performance.
Apply machine learning-based anomaly detection to identify issues proactively and use predictive analytics to forecast system behavior and prevent failures.
Collaborate with cross-functional teams (Engineering, Dev Ops, Security, production support, product) to communicate complex observability concepts to both technical and non-technical stakeholders.
Lead observability initiatives and mentor junior team members on best practices while facilitating design and problem-solving discussions across the organization.
What you need to succeed? Must haves:
5+ years of experience with observability tools (ELK Stack, Dynatrace, Prometheus, Grafana, Open Telemetry, Jaeger, Aternity and Moog or similar).
3+ years of software engineering or infrastructure experience.
Python, Java, GO.
Query languages.
Expert-level knowledge of logging requirements and best practices for enhanced observability.
Demonstrated experience building and optimizing monitoring dashboards.
Proven ability to use observability data to proactively identify and resolve system issues.
Experience using AI tools for anomaly detection and trend analysis.
Alerting.
Linux, Unix.
What's in it for you?
A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, and stock where applicable.
Leaders who support your development through coaching and mentoring opportunities.
Ability to make a difference and lasting impact on system reliability and user experience.
Work in a dynamic, collaborative, progressive, and high-performing team.
The expected salary range for this particular position is $90,000-$140,000, depending on your experience, skills, and registration status, market conditions and business needs.
You have the potential to earn more through RBC's discretionary variable compensation program which gives you an opportunity to increase your total compensation, provided the business meets its performance targets and you meet your individual goals.
RBC's compensation philosophy and principles recognize the importance of a highly qualified global workforce and plays a critical role in attracting, engaging and retaining talent that:
Drives RBC's high-performance culture.
Enables collective achievement of our strategic goals.
Generates sustainable shareholder returns and above market shareholder value
Job Skills
Agile SDLC, AI Observability, Anthropic Claude AI, Database Queries, Data Query Language, Dynatrace Administration, Elastic Stack (ELK), IT Monitoring, Java (Programming Language), Microsoft Copilot, Monitoring Tools, OpenAI, Problem Solving, Python (Programming Language), Site Reliability Engineering, Software Development, SRE Observability, Telemetry Monitoring
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).