Senior Staff Site Reliability Engineer
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-08-23
Listing for:
Cerebras
Full Time
position Listed on 2026-08-23
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
This position is crucial for enhancing the performance and reliability of AI applications. Your expertise will contribute to automated self-service platforms while mentoring junior SREs to create a robust operational culture. Starting with hands-on experience, you will tackle existing production issues and drive change.
Key Responsibilities:
• Architect scalable solutions for multi-datacenter operations
• Build internal tools for streamlined workflow executions
• Define reliability practices for performance measurement
• Guide mid-level SREs during operational incidents
• Analyze and improve metrics for operational efficiency
Requirements:
• 8+ years in SRE or platform engineering roles
• Strong skills in cluster management and CI/CD systems
• Experience in observability with tools like Prometheus
• Proven capability to lead diverse projects
• Ability to communicate technical strategies clearly
Drive the evolution of reliability practices in AI at Cerebras Systems, shaping the future of technology.
#J-18808-Ljbffr
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×