×
Register Here to Apply for Jobs or Post Jobs. X

Remote Infrastructure​/GPU Cluster​/Platform Operations Lead

Remote / Online - Candidates ideally in
Mississauga, Ontario, Canada
Listing for: Bilinguallink
Remote/Work from Home position
Listed on 2026-10-09
Job specializations:
  • IT/Tech
    IT Infrastructure, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Business & Operations
Job Description & How to Apply Below
ELEKS is looking for an Infrastructure/GPU Cluster/Platform Operations Lead in Canada.
Alberta-based candidates are strongly preferred (Calgary or Edmonton). Canada-based candidates will also be considered.

ABOUT CLIENT
Our customer is building a next-generation AI platform that enables organizations to securely develop, govern, and operationalize artificial intelligence while ensuring that sensitive data and organizational knowledge remain fully under their control. The platform combines advanced AI capabilities with enterprise-grade governance, security, and data sovereignty to support mission-critical decision-making.

The solution serves government organizations and enterprise customers operating in highly regulated and security-sensitive environments, where reliability, accountability, and trust are essential. The platform supports intelligent decision-making across strategic planning, workforce intelligence, and organizational operations, helping customers leverage AI without compromising security, compliance, or control over their data.

REQUIREMENTS

8+ years of Infrastructure Engineering or Platform Operations experience

Experience managing GPU clusters for AI workloads

Strong Kubernetes administration skills

Experience with NVIDIA GPU technologies and CUDA ecosystem

Experience with cloud infrastructure (Azure, AWS or GCP)

Knowledge of storage, networking, and high-performance computing environments

Experience implementing Infrastructure as Code (Terraform or similar)

Strong operational leadership skills

Experience supporting AI platform infrastructure

Upper-Intermediate or higher level of English

RESPONSIBILITIES

Lead GPU infrastructure design and operations

Manage Kubernetes-based AI platform environments

Optimize infrastructure for AI training and inference workloads

Define operational standards and reliability practices

Collaborate with AI engineering teams

Implement monitoring, security, and disaster recovery strategies

Lead infrastructure capacity planning

Support technical roadmap and infrastructure evolution

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary