×
Register Here to Apply for Jobs or Post Jobs. X

Observability Engineer

Job in Toronto, Ontario, M5A, Canada
Listing for: Astra North Infoteck Inc.
Full Time position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Job Description & How to Apply Below
Job Description

Observability Engineer – Kubernetes, Prometheus, Grafana & Cloud Monitoring

Role Overview

• Experienced Observability Engineer to join an Enterprise Kubernetes Platform team within a leading financial services organization

• Own and manage the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities

• Build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure using modern observability tools and AI/ML capabilities

Key Responsibilities

• Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents

• Manage observability deployments using Git Ops principles and Infrastructure as Code

• Implement long-term metrics storage solutions using cloud object storage

• Maintain and upgrade observability components across development, QA, UAT, production, and DR environments

• Configure distributed observability architecture spanning multiple datacenters and cloud providers

Metrics & Monitoring

• Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications

• Create Service Monitors and Pod Monitors for automated metrics collection

• Develop intelligent alerting rules with minimal false positives

• Configure multi-cluster metrics federation and aggregation

• Optimize metrics cardinality, storage efficiency, and query performance

Essential Skills

• Kubernetes Observability

• Prometheus

• Grafana

• Thanos

• Loki

• Metrics, Logging, and Tracing

• Alerting and Monitoring

• Git Ops

• Infrastructure as Code

• Cloud Monitoring

• Kubernetes Platform Engineering

• AI/ML-Based Monitoring Solutions

Requirements
Sailpoint
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary