×
Register Here to Apply for Jobs or Post Jobs. X

SLURM HPC Architect​/Administrator

Job in Ottawa, Ontario, Canada
Listing for: Arcadion
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 90000 - 120000 CAD Yearly CAD 90000.00 120000.00 YEAR
Job Description & How to Apply Below

SLURM HPC Architect / Administrator

Location:

Remote (Canada, U.S., or Europe Preferred)
Company:
Cylix Applied Intelligence

Employment Type:

Full-Time or Contract

About the Role

Cylix Applied Intelligence is seeking an experienced SLURM HPC Architect / Administrator to design, deploy, and operate high-performance computing (HPC) clusters supporting AI training, large‑scale inference, scientific computing, and enterprise workloads.

This role will focus on building and managing enterprise‑grade HPC environments powered by GPU and CPU compute clusters, leveraging SLURM as the core workload orchestration and resource scheduling platform.

You will work closely with AI engineers, infrastructure teams, and enterprise clients to deliver scalable, reliable, and high‑performance compute environments across on‑premise, hybrid, and cloud platforms.

Key Responsibilities HPC Cluster Architecture and Design
  • Design and implement SLURM‑based HPC cluster architectures
  • Architect scalable CPU and GPU compute environments
  • Define cluster topology including compute, storage, login, and management nodes
  • Design high‑availability SLURM controller configurations
  • Implement cluster segmentation, partitioning, and resource allocation strategies
SLURM Deployment and Administration
  • Install, configure, and manage SLURM workload manager environments
  • Configure SLURM partitions, queues, QoS policies, and scheduling policies
  • Manage job scheduling optimization and fair‑share policies
  • Implement accounting, usage tracking, and reporting systems
  • Maintain SLURM cluster health, stability, and performance
GPU Cluster and AI Infrastructure Management
  • Configure GPU scheduling and allocation policies
  • Support GPU resource management including:
    • NVIDIA A100, H100, L40, and similar accelerator platforms
    • MIG partitioning and GPU isolation
    • Multi‑tenant GPU resource allocation
  • Optimize cluster performance for AI training and inference workloads
Infrastructure Automation and Operations
  • Automate cluster deployment and configuration using:
    • Ansible, Terraform, or similar tools
    • Shell scripting and Python
  • Implement monitoring, alerting, and performance tracking systems
  • Support cluster lifecycle management, upgrades, and expansion
Storage and File system Integration
  • Integrate HPC clusters with high-performance storage systems including:
    • NFS
    • Lustre
    • BeeGFS
    • GPFS / Spectrum Scale
  • Optimize I/O performance and storage architecture
User and Workload Support
  • Support enterprise and research users with job scheduling and optimization
  • Troubleshoot job failures and performance issues
  • Assist engineering teams in optimizing workloads for HPC environments
Required Qualifications
  • 3+ years experience administering HPC clusters
  • Strong experience with SLURM workload manager
  • Strong Linux system administration experience (Ubuntu, Rocky Linux, RHEL, or similar)
  • Experience with HPC cluster architecture and deployment
  • Experience with shell scripting and automation
  • Experience with:
    • Cluster resource management
    • Multi-node distributed computing environments
    • SSH, networking, and Linux system internals
Preferred Qualifications
  • Experience managing GPU-based HPC clusters
  • Experience supporting AI / ML workloads
  • Experience with NVIDIA GPU platforms and drivers
  • Experience with:
    • CUDA environments
    • NVIDIA MIG configuration
    • GPU scheduling optimization
  • Experience with configuration management tools:
    • Ansible
    • Terraform
    • Puppet or Chef
  • Experience with monitoring tools such as:
    • Prometheus
    • Grafana
    • Node exporter
    • SLURM accounting tools
Nice to Have
  • Experience with large-scale enterprise or cloud HPC environments
  • Experience deploying HPC environments in cloud platforms such as:
    • AWS
    • Azure
    • Private cloud environments
  • Experience with containerized HPC workloads:
    • Docker
    • Singularity / Apptainer
  • Experience integrating SLURM with Kubernetes or hybrid orchestration systems
Exam…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary