Senior Solution Engineer – GPU & AI Infrastructure
Listed on 2026-08-16
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
Senior Solution Engineer – GPU & AI Infrastructure
About Civo:
Civo is a high-performance neocloud provider purpose-built for the demands of modern AI, high-performance computing (HPC), and cloud-native infrastructure. We eliminate legacy cloud overhead to deliver ultra-low-latency compute, bare-metal GPU performance, and streamlined Kubernetes orchestration at scale.
Purpose-designed for AI engineering teams, enterprises, and research institutions, Civo delivers direct access to cutting-edge NVIDIA GPU clusters, high-speed fabrics, and parallel storage systems required to train, fine-tune, and deploy foundation models efficiently. We combine high-density infrastructure with predictable pricing and maximum compute throughput, empowering organizations to scale AI workloads without the complexity or cost bloat of traditional hyperscalers.
About the Role:
As a Senior Solution Engineer – GPU & AI Infrastructure, you will serve as the primary technical architect for Civo’s large-scale AI and high-performance computing (HPC) customer initiatives. You will be responsible for designing state-of-the-art NVIDIA GPU clusters tailored for training and inferencing massive foundation models.
In this role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare-metal and Kubernetes-based orchestrations across cutting-edge NVIDIA Blackwell architectures (e.g., B300 and GB300NV
L) using ultra-low-latency Infini Band and high-speed RoCE networking fabrics.
Responsibilities:
Solution Design & Architecture
- System Design Documents:
Author comprehensive High-Level Design (HLD) and Low-Level Design (LLD) documentation for enterprise-scale GPU supercomputing clusters. - Bill of Materials (BOM):
Generate detailed BOMs covering compute nodes, NVLink switches, network fabrics, transceivers/cabling, liquid/air cooling requirements, power distribution, and high-performance storage. - GPU Cluster Topology:
Architect scale-up (NVLink/NVSwitch) and scale-out network topologies (Fat-Tree, Rail-Optimized) for NVIDIA Blackwell platforms, specifically B300 and GB300NV
L rack-scale architectures. - Fabric & Networking Engineering:
Design high-throughput, low-latency networking architectures utilizing both Infini Band (e.g., NDR/X800) and RoCE / RoCEv2 (e.g., NVIDIA Spectrum-X / Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). - Multi-Tenant & Deployment Models:
Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native / Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, Kube Flow). - Storage Integration:
Architect high-bandwidth parallel storage solutions utilizing GPUDirect Storage (GDS) and enterprise AI file systems (e.g., VAST Data).
Technical Sales Support & Customer Engagement
- Partner with Civo’s sales and commercial teams as the technical lead for high-value AI infrastructure opportunities.
- Engage directly with customer CTOs, Chief AI Officers, infrastructure leads, and ML engineers to evaluate technical requirements, compute sizing, and fabric choices.
- Lead deep-dive architectural workshops and technical presentations on Civo's bare-metal GPU and managed Kubernetes offerings.
- Produce precise technical proposals and lead responses to complex RFPs/RFIs regarding AI infrastructure.
Proof-of-Concept (PoC) & Benchmarking
- Architect and oversee Proof-of-Concept (PoC) deployments to validate real-world performance for customer workloads.
- Benchmark cluster performance using industry-standard tools (NCCL tests, GPUDirect RDMA latency/bandwidth, MLPerf, Megatron-LM benchmarks).
- Address network congestion, fabric routing, and thermal/power optimization during validation phases.
Product & Ecosystem Collaboration
- Serve as the bridge between enterprise AI clients, hardware vendors (NVIDIA, network OEMs), and Civo’s internal platform engineering team.
- Provide continuous feedback to product teams on…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: