Principal Solution Architect – AI Infrastructure and Cloud
Listed on 2026-09-12
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
IREN is a vertically integrated AI Cloud provider, delivering large-scale data centers and GPU clusters for AI training and inference. IREN’s platform is underpinned by its expansive portfolio of grid-connected land and power in renewable-rich regions across North America, Europe and APAC.
With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever-evolving applications of high-performance compute. We believe that human progress is invaluable, but it should be done in the right way – responsibly, sustainably and having a positive impact on the communities we operate in.
As a Principal Solution Architect, you will sit at the tip of the spear for IREN's AI Cloud platform. You will partner directly with client CTOs, lead ML engineers, and infrastructure architects to design, benchmark, and optimize massively distributed training and inference clusters. You will convert complex customer requirements into optimized hardware/software configurations while creating reference architectures that set industry benchmarks.
BasicQualifications
- Experience:
8–10+ years as a Solution Architect, Systems Engineer, or Infrastructure Architect handling enterprise cloud, HPC, or AI environments. - AI & Hardware Mastery:
Deep hands-on experience with high-density GPU nodes (NVIDIA HGX/DGX/Blackwell), Linux bare-metal virtualization, and RDMA networking (Infini Band/RoCE v2). - Orchestration & Storage:
Proficiency with Kubernetes, container networking (CNI), and parallel file systems (Lustre, VAST, Ceph). - Customer-Facing Execution:
Proven ability to lead complex technical workshops, present to C-suite executives, and resolve high-stakes production bottlenecks.
- Practical experience tuning NCCL communication, GPUDirect Storage (GDS), or custom Infini Band routing fabrics.
- Contributions to open-source cloud-native or AI orchestration projects (e.g., CNCF, Ray, Slurm, k0rdent).
- B.S. or M.S. in Computer Science, Computer Engineering, or quantitative discipline.
- Serve as the trusted lead technical architect for tier-1 enterprise customers migrating or scaling AI workloads on IREN’s infrastructure.
- Design, document, and publish end-to-end reference architectures covering GPU clusters, Infini Band/RoCE network topologies, Ceph/Lustre storage fabrics, and Kubernetes multi-tenancy.
- Lead hands‑on proof-of-concept deployments, conducting cluster benchmarking (e.g., NCCL-tests, Linpack, MLPerf, IO500) to validate customer SLA requirements.
- Partner with client engineering teams to profile and optimize distributed training jobs (Megatron‑LM, Deep Speed, PyTorch FSDP) and inference pipelines (vLLM, TensorRT‑LLM).
- Translate customer pain points, telemetry, and edge cases into structured platform requirements for IREN’s core Product and Operations teams.
- Actual compensation will be determined based on factors such as experience, qualifications, and market data for the region.
- Total Compensation package may be inclusive of short-term and long-term incentives.
- Relocation (as applicable and based on successful candidate circumstances)
- Medical, dental, and vision insurance coverage – 100% company paid for employees, 75% company paid coverage for dependents
- Company-paid life and disability insurance
- Voluntary life, critical illness, and accident coverage available
- Health Savings Accounts (HSA) – when combined with the High-Deductible Health Plan
- Employee Assistance Program and wellness resources
- 401(k) retirement plan with company match
- Financial…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).