×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Waterloo, Kitchener, Ontario, Canada
Listing for: Magnet Forensics
Full Time position
Listed on 2026-08-25
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Cybersecurity
Job Description & How to Apply Below
Location: Waterloo

Who We Are;
What We Do;
Where We’re Going
Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices. We are continually innovating so our customers can deploy advanced and effective tools to protect their companies, communities, and countries. Serving thousands of customers globally, our solutions are playing a crucial role in modernizing digital investigations, helping investigators fight crime, protect assets, and guard national security.

With employees based around the world, Magnet Forensics has been expanding our global presence. As a part of Magnet Forensics, you can expect to make a difference in the world, no matter what role you play. You’ll be supported through learning and development, not to mention an incredible team with unbelievable talent and integrity. If you think you would be the right person to join our team working towards this goal, we would love to hear from you!

Role Overview

We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly available SaaS platform, a production Kubernetes environment serving law enforcement and government customers globally. This role requires deep AWS expertise, infrastructure-as-code discipline, and CI/CD best practices. You'll work closely with Application, Platform, and Security teams to drive secure-by-design architectures and improve automation and reliability across our cloud environments.

You'll ship infrastructure as code, respond to production incidents with discipline, and drive platform modernization through deliberate roadmap execution. As part of the SaaS-Ops team, you’ll work in a high-performing environment where members take ownership of outcomes and operate with a strong sense of trust and autonomy. You’ll identify challenges, contribute to solutions, raise concerns proactively, support improvements, and navigate situations requiring timely decision-making.

If you’re looking for your next challenge where infrastructure quality directly impacts real‑world outcomes, this role could be a great fit!
Note:
This role includes participation in an on-call rotation.

What You’ll Do

  • Own and operate production Kubernetes clusters (Amazon EKS) including upgrades, scaling, security hardening, and cluster lifecycle management;
  • Design, implement, and maintain infrastructure-as-code using Terraform; contribute to shared module libraries and enforce IaC standards across the team;

  • Manage and evolve Helm chart definitions and ArgoCD Git Ops workflows for multi-region SaaS deployments;

  • Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log pipelines. Act to eliminate noise and surface signal;

  • Contribute to pipeline reliability: identify flaky stages, reduce build times, improve developer experience across CI/CD pipelines;

  • Remediate security vulnerabilities (CVEs) in container images and infrastructure components; participate in compliance work including FedRAMP support activities;

  • Develop and maintain runbooks, change management procedures, and operational documentation;

  • Ensure alignment with internal policies and frameworks such as ISO 27001, SOC2, and NIST;

  • Contribute to AI-assisted tooling and automation (, Claude-based Terraform agents, automated triage tools) as part of the team's operational efficiency roadmap;

  • Participate in on-call incident response rotation; lead or support incident command during active production incidents including root cause analysis and post-incident review.

  • What We’re Looking For

  • 5+ years of industry experience with a trajectory that demonstrates growing depth in cloud infrastructure and SRE practices;

  • Managed production Kubernetes environments at scale: not just deployed workloads, but owned cluster health, upgrades, and failure modes;

  • Responded to production incidents in high-stakes environments where downtime has real consequences;

  • Written and maintained Terraform at the module level, not just as a consumer: understands state,…

  • Position Requirements
    10+ Years work experience
    Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
    To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary