Staff Site Reliability Engineer (Copy
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-08-12
Listing for:
Socket.dev
Full Time
position Listed on 2026-08-12
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
The Opportunity
We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.
WhatYou’ll Do
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
- Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
- Partner with engineering teams to optimize performance, cost, and reliability of backend services.
- Drive incident response, post-mortem analysis, and long-term remediation efforts.
- Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
- Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
- Lead department wide compliance (PCI, SOC) initiatives.
- 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or Dev Ops roles.
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
- Experience designing and managing high-throughput, distributed systems.
- Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
- Excellent communication skills and a demonstrated ability to mentor engineers.
- Experience working in a high-velocity, customer-focused environment.
- Familiarity with functional programming paradigms (e.g., Clojure/Clojure Script).
- Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
- Experience implementing security and compliance best practices in the cloud.
This role is available throughout Canada.
Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits.
Our recruiting team will provide a specific salary range based on location and years of experience.
#J-18808-LjbffrNote that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×