Principal Solution Specialist, Core Services
Listed on 2026-08-26
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, Network Engineer
Core Weave is The Essential Cloud for AI™. Built for pioneers by pioneers, Core Weave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, Core Weave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, Core Weave became a publicly traded company (Nasdaq: CRWV) in March 2025.
Learn more at
Core Weave’s core infrastructure (bare-metal GPU compute, high-performance networking, and the Kubernetes platform that ties them together) is constantly evolving, and bringing each new capability to market is where this role lives. As a Principal Solution Specialist for Core Services, you identify the new market opportunities where compute performance, network fabric architecture, and platform reliability unlock adoption, and you drive the first wins for Core Weave’s newest core capabilities with the customers and industries that demand them earliest.
You turn field insight into direct input on the compute and networking roadmap, and you make the broader sales and solution architecture teams fluent in the value these services create for new customers.
About the role: As a Principal Solution Specialist for Core Services, you operate at the front edge of Core Weave’s go-to-market motion for its foundational infrastructure. Your job is to create the motion, not just run it: you take new and evolving core offerings spanning our Compute (bare-metal and virtualized GPU instances) and Networking (high-speed fabrics and interconnects for tight GPU-to-GPU communication) into new accounts and new industries, prove their value with the most demanding early customers, and build the repeatable playbooks sales and solution architects use to scale.
You are the field’s voice into engineering, translating what early adopters need around cluster topology, RDMA networking, and scheduling into the priorities that shape Core Weave’s compute and networking roadmap.
In this role, you will:
- Own the commercial and technical strategy for net new customer wins in core AI infrastructure, where GPU compute performance, cluster networking, and platform reliability are the primary buying triggers.
- Drive new business opportunities where compute bottlenecks, network bandwidth limitations, or bare-metal performance requirements are the barrier to scaling AI training and inference workloads on Core Weave.
- Translate customer requirements around GPU cluster topology, RDMA networking (Infini Band, RoCE), and Kubernetes scheduling into specific product feedback that shapes Core Weave’s compute and networking roadmap.
- Develop deal structures, technical playbooks, and benchmark narratives that help sales and SA teams accelerate compute-heavy and networking-sensitive opportunities.
- Engage directly with enterprise and research buyers as the authoritative voice on GPU cluster architecture, network fabric design, and the performance tradeoffs between instance types and topologies.
- Design the commercial framework for large‑scale GPU reservation deals, including MFU modeling, cluster sizing, and network bandwidth commitments that support large enterprise closings.
- Partner with capacity, infrastructure, and networking teams to maintain a competitive edge on compute density, interconnect performance, and platform reliability across active and prospective customer deployments.
- 10+ years of experience in HPC, data center infrastructure, or GPU cluster engineering, with a track record of applying that expertise to drive customer outcomes and revenue.
- 5+ years working with large-scale GPU clusters (NVIDIA A100/H100/H200 or AMD MI300X), RDMA networking, and distributed training frameworks in a customer‑facing or deal‑shaping capacity.
- Deep working knowledge of high‑speed network fabrics (Infini Band, RoCE, and NVLink) and how interconnect topology, bandwidth, and latency impact distributed training and inference at scale.
- Experience deploying and tuning large‑scale Kubernetes clusters for GPU workloads (NVIDIA GPU…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).