×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer (SRE), Cloud Operations

Job in Toronto, Ontario, C6A, Canada
Listing for: United States Digital Space LLC
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 110000 - 140000 CAD Yearly CAD 110000.00 140000.00 YEAR
Job Description & How to Apply Below

What is the opportunity?

Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms — from Open Shift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices.

If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations — this is the role.

What will you do?
  • Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/Open Shift, ECE, Confluent Kafka).
  • Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
  • Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
  • Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
  • Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., Service Now, Prometheus) that support operational excellence.
  • Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
  • Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
  • Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and Pager Duty.
What do you need to succeed?

Must-have

  • 5+ years of hands‑on experience in Site Reliability Engineering, Dev Ops, or infrastructure operations.
  • Strong working knowledge of Kubernetes/Open Shift administration and troubleshooting in enterprise environments.
  • Hands‑on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands‑on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on‑call rotations, and post‑incident review practices.
  • Familiarity with capacity planning, threshold‑based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).

Nice-to‑have

  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
  • Hands‑on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi‑cloud environments.
  • Experience with GPU/compute infrastructure for ML inference workloads
What’s in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high‑performing team
  • Flexible work/life balance…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary