More jobs:
Expert Site Reliability Engineer at IBM
Job in
Toronto, Ontario, C6A, Canada
Listed on 2026-07-21
Listing for:
IBM Computing
Full Time
position Listed on 2026-07-21
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support, Systems Engineer
Job Description & How to Apply Below
In this hands-on role, you will focus on analyzing systemic failures and instituting proactive improvements. Your time will be split between development work and mentoring, ensuring teams are fully equipped to handle incidents effectively. Collaborate with a global group of engineers dedicated to excellence in service delivery.
Key Responsibilities:
• Investigate and improve incident recurrence prevention strategies
• Oversee integration of incident management tools like Rootly and Pager Duty
• Implement and maintain SLO/SLA frameworks for reliability
• Facilitate training programs and lead post-mortem discussions
• Edit documentation ensuring quality in incident reporting
Requirements:
• 10+ years in site reliability or incident management
• Experienced with cloud services like AWS, GCP, or Azure
• Deep knowledge of incident management tools
• Strong grasp of distributed systems complexity
• Experience with Kubernetes and CI/CD pipeline understanding
Join us to enhance reliability in a diverse, innovative environment at IBM.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×