Vice President of Infrastructure & Deployment; Remote/Hybrid
The Woodlands, Montgomery County, Texas, USA
Listed on 2026-07-22
-
IT/Tech
Systems Engineer, IT Infrastructure, Network Engineer
Vice President of Infrastructure & Deployment
Pay: $ - $
Equity: 0.25-0.5%
Location & TravelRemote or hybrid, with regular travel to data center markets, customer deployment sites, vendor locations, and company operating sessions expected.
About AcasiaAcasia builds, deploys, and operates high-performance GPU infrastructure for enterprise AI workloads. Our customers rely on Acasia to deliver production-grade GPU environments that are performant, reliable, scalable, and supportable in real-world data center conditions.
Role SummaryThe Vice President of Infrastructure & Deployment owns the successful delivery, implementation, commissioning, and operational readiness of Acasia's AI infrastructure across customer and data center environments.
This executive owns the complete delivery lifecycle—from customer handoff after contract execution through deployment planning, installation, networking, cluster bring‑up, validation, production acceptance, and ongoing optimization—ensuring every GPU cluster is delivered safely, on schedule, on budget, and to Acasia's standards before entering production.
The VP will build and lead Acasia's Infrastructure Delivery organization, managing field engineering teams that implement large‑scale GPU clusters across multiple data centers and customer locations. They will establish deployment methodologies, technical standards, playbooks, commissioning procedures, and quality controls while driving continuous improvement in deployment speed, consistency, and customer experience. They will work hand‑in‑hand with customers, Sales, Customer Success, Engineering, Product, Operations, Procurement, OEM partners, networking vendors, and data center operators.
MissionBuild the industry's fastest, most reliable GPU infrastructure deployment organization—taking customer environments from signed contract to production in weeks instead of months.
Key Responsibilities Infrastructure Architecture & Technical Leadership- Own infrastructure architecture standards for GPU clusters across all Acasia deployments.
- Define reference architectures for rack layouts, power distribution, cooling, networking, storage, monitoring, telemetry, and remote management.
- Review customer infrastructure designs for scalability, resilience, serviceability, deployment efficiency, and lifecycle management.
- Identify technical and deployment risks before implementation begins.
- Establish repeatable infrastructure standards that simplify deployment while improving reliability and operational excellence.
- Partner with Engineering, Product, Security, Sales, Customer Success, Procurement, OEMs, and customers to ensure infrastructure designs meet both technical and commercial objectives.
- Own deployment and validation of high‑performance AI networking environments.
- Lead implementation of Infini Band, Ethernet, RoCE, RDMA, NVLink, NVSwitch, GPUDirect, and NVIDIA networking technologies.
- Oversee spine‑leaf architecture deployment, east‑west networking, switch configuration, storage networking, and cluster connectivity.
- Establish standards for network validation, benchmarking, congestion analysis, latency optimization, and fabric performance tuning.
- Validate cluster communication performance using NCCL, MPI, RDMA, and other AI infrastructure benchmarking tools.
- Partner with NVIDIA and networking OEMs to optimize production environments.
- Define production acceptance standards for every customer deployment.
- Own Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), cluster commissioning, and customer handoff.
- Establish standards for hardware acceptance, BIOS and firmware consistency, GPU validation, driver installation, CUDA readiness, storage validation, telemetry, monitoring, and observability.
- Lead burn‑in procedures, stress testing, soak testing, redundancy validation, performance baselining, and workload validation before production acceptance.
- Ensure every customer environment meets Acasia's technical standards before entering production.
- Build and manage deployment schedules for large‑scale infrastructure implementations.
- Coordinate execution across…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).