×
Register Here to Apply for Jobs or Post Jobs. X

Head of Site Reliability Engineering

Job in Toronto, Ontario, C6A, Canada
Listing for: Shakudo
Full Time position
Listed on 2025-12-21
Job specializations:
  • IT/Tech
    Cloud Computing, Systems Engineer, SRE/Site Reliability, IT Support
Salary/Wage Range or Industry Benchmark: 120000 - 160000 CAD Yearly CAD 120000.00 160000.00 YEAR
Job Description & How to Apply Below

About the Job & Shakudo

At Shakudo, we are building the world’s first operating system for data and AI. We use the term operating system in the truest sense of the word. Like iOS, Windows and Linux, Shakudo’s end-to-end OS offers ever-evolving, automatically operated, best-of-breed open-source components tailored to each business's unique needs.

The Role

We are hiring a Head of Site Reliability Engineering to lead the reliability, availability, and performance strategy of our platform. This role is ideal for someone who thrives on solving infrastructure challenges, scaling cloud-native systems, and building high-performance teams.

You will work cross-functionally with engineering, product, and customer success to make Shakudo’s platform rock-solid and resilient for our customers around the world.

What You’ll Do
  • Build and lead the SRE function at Shakudo, setting goals, technical direction, and driving team culture
  • Own uptime, reliability, and incident response for our platform
  • Architect scalable infrastructure using Kubernetes, cloud-native tooling, and automation frameworks
  • Lead the design of observability, monitoring, and alerting systems to proactively detect and prevent issues
  • Create and enforce best practices for CI/CD, disaster recovery, and service-level objectives (SLOs)
  • Partner closely with engineering and product to ensure new features are reliable and production-ready
  • Mentor engineers and help instill a culture of operational excellence
What We're Looking For
  • 8+ years of experience in infrastructure, Dev Ops, or SRE roles with increasing responsibility
  • Proven experience scaling distributed systems in a high-availability, production environment
  • Expertise with Kubernetes, Terraform, containerization, and at least one major cloud provider (AWS preferred)
  • Strong knowledge of system design, networking, and reliability principles
  • Experience with observability tools (e.g., Prometheus, Grafana, Datadog) and incident response practices
  • Strong leadership and communication skills, with a hands-on, collaborative approach
Nice to Have
  • Experience supporting data pipelines, ML workloads, or complex orchestration systems
  • Familiarity with the data/ML tooling ecosystem (e.g., Airflow, dbt, Spark, Dremio, etc.)
  • Previous experience in a startup or high-growth environment

Shakudo is an equal opportunity employer and encourages candidates of all backgrounds to apply. We foster diversity and inclusivity and welcome applications from a broad range of backgrounds and experiences.

#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)

Job Posting Language
Employment Category
Education (minimum level)
Filters
Education Level
Experience Level (years)
Posted in last:
Salary