×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Specialist, IT Operations

Remote / Online - Candidates ideally in
Province de Québec, Canada
Listing for: SherWeb
Full Time, Remote/Work from Home position
Listed on 2026-08-29
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Support, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Location :
Work from Home - Province of Quebec

The Site Reliability Specialist on the IT Operations team contributes to the reliability, availability, performance, and resilience of Sherweb's platforms and services.

This is a highly technical individual contributor role that applies Site Reliability Engineering (SRE) principles to production environments. The role combines systems administration, software development, automation, observability, and operational excellence to improve service reliability, reduce operational toil, and increase platform scalability.

Working closely with Infrastructure, Development, Dev Ops, Platform, Security, and Product teams, the Site Reliability Specialist helps ensure production systems remain stable, supportable, and continuously improving through engineering and automation practices.

Here's how you will contribute to the success of the company

· Develop, maintain, and improve scripts, automation workflows, and operational tooling to reduce manual intervention, improve reliability, and lower operational toil.

· Apply SRE principles to improve the reliability, availability, performance, and resilience of Sherweb’s hosted platforms and production services.

· Implement and support reliability standards, service level objectives (SLOs), service level indicators (SLIs), and operational practices established for platforms and services.

· Use a developer mindset to transform repetitive operational tasks into scalable, reusable, documented, and supportable automation.

· Provide advanced operational support and resolve incidents affecting production services while collaborating closely with SRE, Infrastructure, Development, Dev Ops, and Product teams.

· Investigate recurring issues and perform root cause analysis to identify short-term corrective actions and long-term reliability improvements.

· Build, support, and maintain production systems and hosted service technologies while following operational procedures, security best practices, and compliance requirements.

· Improve monitoring, alerting, logging, metrics, and operational visibility to help the team detect issues earlier, understand system behavior, and prevent incidents.

· Contribute to improving end-to-end observability and system understanding through metrics, logs, traces, telemetry, and operational diagnostics.

· Contribute to observability-as-code, infrastructure-as-code, configuration-as-code, and automation practices where applicable.

· Explore and leverage Azure AI Foundry, Power Automate, and AI agent capabilities to improve operational efficiency, automate repetitive workflows, and accelerate incident response or service reliability improvements.

· Participate in platform lifecycle activities, deployments, maintenance windows, migrations, and continuous service improvement initiatives.

· Collaborate with developers, architects, subject matter experts, Dev Ops, and infrastructure teams to support implementation, optimization, troubleshooting, and operational readiness of services.

· Create and maintain operational documentation, including SOPs, runbooks, troubleshooting guides, automation documentation, maintenance procedures, and knowledge-sharing materials.

· Track, organize, and manage incidents, requests, and service tickets to respect SLAs and ensure clear communication through the ITSM process.

· Participate in rotational on-call duty and perform maintenance work outside normal business hours when required.

· Carry out all other related tasks per the job’s evolution and departmental needs.

Here's what you need to have and master to get the job

Education

· College or university degree in computer science, software development, information technology, engineering, or a combination of equivalent training and experience.

Experience

· 3 to 5 years of experience in systems administration, IT operations, infrastructure support, software development, Dev Ops, automation, or a similar technical role.

· Experience supporting production systems in business-critical and customer-facing environments.

· Proven experience improving operational efficiency through automation and engineering practices.

Core Skills

· Strong scripting or…

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary