Site Reliability Engineer – APM, Dynatrace, Observability
Listed on 2026-07-21
-
IT/Tech
SRE/Site Reliability
Site Reliability Engineer – APM, Dynatrace, Observability Location
- Toronto, ON
- Hybrid – 2 days in office per week
Strong Day 1 expertise in Dynatrace, including:
- DQL
- Gen3 Dashboards
- Traces / Grail
- Active Gate and Plugins
- SRG / Workflow Development
- Biz Events
Hands‑on experience with APM and observability platforms (Dynatrace or equivalent), with the ability to instrument, analyze, and troubleshoot distributed applications.
Deep troubleshooting expertise using observability signals (Metrics, Events, Logs, Traces – MELT) to identify root causes across complex, multi‑layer end‑to‑end environments.
Strong foundation in Observability fundamentals (MELT).
Expert‑level dashboard design, including UI/UX best practices.
Extensive experience troubleshooting performance and non‑functional issues.
Strong expertise in AWS Observability, including:
- Cloud Watch
- Application Signals
- Metrics, Logs, and Traces
- AWS Lambda
- API Gateway
Development experience with:
- Python
- Ansible
- AWS Lambda
- Amazon ECS
- Azure Functions
Understanding of AI‑based system fundamentals, including how AI systems are built and monitored.
#J-18808-LjbffrTo Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: