More jobs:
Remote Infrastructure/GPU Cluster/Platform Operations Lead
Remote / Online - Candidates ideally in
Mississauga, Ontario, Canada
Listed on 2026-10-09
Mississauga, Ontario, Canada
Listing for:
Bilinguallink
Remote/Work from Home
position Listed on 2026-10-09
Job specializations:
-
IT/Tech
IT Infrastructure, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Business & Operations
Job Description & How to Apply Below
Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.
ABOUT CLIENT
Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making.
The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.
REQUIREMENTS
8+ years of Infrastructure Engineering or Platform Operations experience
Experience managing GPU clusters for AI workloads
Strong Kubernetes administration skills
Experience with NVIDIA GPU technologies and CUDA ecosystem
Experience with cloud infrastructure (Azure, AWS or GCP)
Knowledge of storage, networking, and high-performance computing environments
Experience implementing Infrastructure as Code (Terraform or similar)
Strong operational leadership skills
Experience supporting AI platform infrastructure
Upper-Intermediate or higher level of English
RESPONSIBILITIES
Lead GPU infrastructure design and operations
Manage Kubernetes-based AI platform environments
Optimize infrastructure for AI training and inference workloads
Define operational standards and reliability practices
Collaborate with AI engineering teams
Implement monitoring, security, and disaster recovery strategies
Lead infrastructure capacity planning
Support technical roadmap and infrastructure evolution
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×