Site Reliability Specialist
Job in
Montreal, Montréal, Province de Québec, Canada
Listing for:
Ubisoft
Full Time
position
Listed on 2026-07-27
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Location: MontrealJob Description
As a Site Reliability Specialist at Ubisoft Montréal, you will join the IT Games and Studios team and help improve the availability, reliability, and performance of critical platforms and services that support game development across Ubisoft.
You’ll collaborate with developers, cloud specialists, and infrastructure teams to build resilient solutions, strengthen observability practices, and support operational excellence across production environments. Through automation, continuous improvement, and reliability-focused initiatives, you’ll help ensure stable and efficient services for teams across the organisation.
What you’ll do
Collaborate with service teams to define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs)Design and implement automation solutions that improve operational efficiency and service reliabilityDocument technical solutions and support their integration across internal platforms and servicesAlign technical implementations with established quality standards, engineering practices, and IT guidelinesPartner with development, infrastructure, and platform teams to improve operational consistency and system reliabilitySupport observability practices, including monitoring, logging, alerting, and incident managementContribute to root cause analysis and continuous improvement initiatives following service incidentsOptimise deployment workflows and operational processes through automation and infrastructure improvementsSupport the maintenance and evolution of cloud-based environments and servicesShare knowledge and contribute to reliability-focused initiatives across teamsQualifications
What you bring to the team
Experience with infrastructure engineering, automation, and Dev Ops practicesKnowledge of Git Lab and Git Lab CI/CD for deployment and automation workflowsProficiency with scripting and programming languages such as Python, Bash, and GoExperience using Terraform, Infrastructure as Code (IaC) practices, and Kubernetes (K8s) in public cloud environments such as Amazon Web Services (AWS) or Microsoft AzureFamiliarity with configuration management tools such as Ansible or ChefExperience with observability and monitoring platforms such as Prometheus and GrafanaInterest in AI-assisted engineering tools such as Git Hub Copilot, Claude Code, or similar solutions that support development, automation, and operational efficiencyAbility to collaborate effectively with technical and non-technical partners while supporting problem-solving and continuous improvement initiatives
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here: