×
Register Here to Apply for Jobs or Post Jobs. X

Cluster Site Reliability Engineer

Job in Regina, Saskatchewan, S4M, Canada
Listing for: iFrame Corporation
Full Time position
Listed on 2026-08-25
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 210000 - 340000 CAD Yearly CAD 210000.00 340000.00 YEAR
Job Description & How to Apply Below

Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate Infini Band fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware.

Type Full-time
· IC4-IC6

Stack Linux Infini Band (NDR / XDR) NCCL / RCCL Kubernetes (host-level) Terraform Prometheus / Grafana Go or Python

Hiring manager replies within 5 business days.

The team

About the team

Cluster SRE is six engineers across the seven regions. Each region has a primary and a secondary; you will be one of those for your region. The team coordinates daily, deploys weekly, and rotates a global pager.

Reports to the head of cluster engineering. Primary on a single region; rotates secondary for one neighboring region.

What you'll do
  • 01 Bring up new B200 / B300 / MI300X racks: cabling, ToR config, NCCL/RCCL all-reduce validation, MFU baseline tests.
  • 02 Drive Infini Band fabric to spec - NDR / XDR depending on the rack - and chase residual bit-error budget down to zero.
  • 03 Run capacity planning across seven regions: forecast demand, model power and thermal headroom, work with procurement on lead times.
  • 04 Own the regional incident response. P1 incidents page within fifteen minutes; resolution target is four hours.
  • 05 Build and maintain the bring-up runbook so the second hire after you can do their first rack solo.
  • 06 Carry the global pager about one week per six, alongside runtime and customer engineering.

The bar

What we're looking for

Five-plus years operating large compute clusters — supercomputing centers, hyperscaler infra, or HPC at a national lab count.

Deep Infini Band and Ethernet RoCE experience: subnet manager tuning, fabric debugging, lossless networking.

Comfort writing Go or Python for tooling. We are not strict about which.

Calm under load. You will be the named person on a $50M-ARR account when something goes wrong.

Nice to have, not required

Experience with NVIDIA Bright / Base Command Manager.

Bare-metal provisioning systems:
Tinkerbell, MAAS, Razor, or in-house equivalents.

Compensation In writing, like everything else

We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.

Base

$210,000 – $340,000 USD (US Cologix regions) / equivalent in CA.

Equity

Meaningful early-stage equity, refreshed on tenure milestones.

Notes

On-site pay differential at Cologix regions outside SF / NYC / Bay Area is +5-10% to compensate for travel.

Equal opportunity

We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU — bring it up on the manager call.

#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary