Technical Program Manager, Deployments
Listed on 2026-07-27
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About This Role:Crusoe is the world's first vertically integrated, sustainable AI cloud. We build and operate GPU infrastructure powered by clean energy, from data center design through IaaS products to managed inference at scale, enabling AI-native companies to run demanding workloads without compromising on sustainability or reliability. Crusoe Cloud is 1,400 people and growing. The TPM frameworks are still being built, which means there is a real opportunity to shape how the function operates rather than inherit how it already works.
We are hiring a Staff TPM who will own deployment programs. You will be helping us build out new sites, capacity expansion, or help deploy modular (Spark) data centers. Our TPMs define and drive the entire deployment engagement model from chip vendor engagement through first customer cluster delivery.
If you have spent your career running infrastructure deployment programs at a hyperscaler or a leading neocloud, understand GPU architecture, and have shaped deployment frameworks rather than just operated within them, this is the role for you.
What You'll Be Working On:New Deployments
- Own the infrastructure deployment for new sites or site expansions end-to-end: chip vendor and OEM dependencies, architecture updates, cloud foundations work, commissioning gate framework definition, and first customer cluster delivery as the success metric.
- Lead Deployment Phase 0 on the Cloud TPM side: define firmware version targets and DOCA targets before kickoff, set commissioning gate criteria, and define the DRI matrix.
- Manage compounding cross-SKU dependencies where active production programs (e.g. B200/GB300/VR) are running in parallel with capacity expansion projects, and prevent them from competing for the same Engineering pool without a plan.
- Owns real-time execution dashboards; delivers crisp, data-driven executive updates that surface decision elements without requiring follow-up
- Governs cross-organizational dependencies without waiting for escalation authority
- Coach more junior TPMs on technical depth, risk identification, and executive communication.
- Actively drives AI tool integration across their programs; identifies where AI materially improves program tracking, risk detection, and executive communication
Technical Foundation
- Deep, working fluency with GPU architecture across SKU generations, firmware lifecycle (DOCA, driver stacks, BIOS/BMC), compute orchestration, SDN, storage, networking (leaf-spine topology, ZTP, fabric commissioning), and monitoring/observability
- Direct hardware partner engagement: personal ownership of NVIDIA or OEM certification and validation timelines, not coordination feeding into someone else's relationship.
- Active daily use of AI tools to drive program-level outcomes: risk detection, dependency mapping, data analysis, and executive communication, not just personal productivity.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).