Member of Technical Staff - AI Infrastructure
Listed on 2026-08-08
-
Software Development
DevOps, Cloud Engineer - Software, Unix/Linux
Ready to take the next step in your career?
Join a fast-growing AI compute platform building the next generation of agentic infrastructure for GPU-intensive workloads. Operating across the full technology stack, the organisation combines high-performance hardware with intelligent software to automate workload placement, scaling, recovery, and orchestration for AI infrastructure at scale.
This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help build the software platforms powering the next generation of AI infrastructure, working on orchestration intelligence, inference gateways, and agentic operations tooling that underpin a live production environment serving real customers.
Responsibilities:- You will be building and evolving the core AI infrastructure software stack, orchestration, scheduling, cluster management, and the agentic operations layer that automates placement, healing, and recovery
- You will be contributing to the inference platform, including the token gateway, microVM pooling, and model serving infrastructure built for high utilisation and low cold-start latency
- You will be working across bare metal, Slurm, Kubernetes, and Infini Band environments, writing software that abstracts the complexity for customers without hiding it from the engineers who need to debug it
- You will be building Infrastructure as Code tooling that automates deployment at scale and tracks changes across a fleet of 10,000+ servers
- You will be integrating with NVIDIA inference microservices, custom model pipelines, and customer-specific workloads — writing software that makes complex GPU infrastructure feel simple to consume
- You will be contributing to the SRE tooling layer; monitoring, continuous optimisation, GPU utilisation balancing, and automated remediation
- Strong systems software background; you will be comfortable at the intersection of distributed systems, infrastructure automation, and low-level performance
- Experience building orchestration, scheduling, or cluster management software at scale;
Kubernetes, Slurm, or equivalent - Solid understanding of GPU infrastructure; you will not need to crimp cables, but you will need to understand what Infini Band, NVLink, and NCCL mean for the software you write
- Production engineering mindset; you will have shipped software that runs in anger on real infrastructure, not just in staging
- Strong Go, Python, or Rust; comfort with Linux systems programming underneath it all
Skills:
- Experience building agentic or autonomous systems that operate infrastructure without human-in-the-loop
- Background in inference serving, vLLM, TensorRT-LLM, NVIDIA NIM, or equivalent
- Exposure to bare metal provisioning, IPMI/BMC management, or large-scale server fleet tooling
- Familiarity with hypervisor or microVM technology (Firecracker, Cloud Hypervisor, or similar)
- Early-stage equity
- Direct access to leadership and genuine technical ownership
- $250,000 Base salary
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).