Virtualization & Orchestration Engineer
Listed on 2026-07-24
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Infrastructure
Virtualization & Orchestration Engineer
Location: Hybrid | Bellevue, WA Area
Titles: Intermediate, Senior and Staff (multiple roles available)
About the Opportunity
A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.
Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern automation and tooling to build infrastructure capable of supporting the industry's most demanding AI workloads.
We're seeking Virtualization & Orchestration Engineers to build the platform layer that enables customers to reliably consume GPU compute s team is responsible for designing and operating the virtualization, Kubernetes, provisioning, and orchestration systems that power large-scale AI workloads across next-generation data center infrastructure.
The OpportunityThis is a foundational engineering role within the company's largest infrastructure engineering organization. You'll help design and build the systems that make GPU capacity available, scalable, secure, and reliable across a multi-tenant AI cloud platform.
You'll work at the intersection of virtualization, Kubernetes, distributed systems, GPU infrastructure, and high-performance computing—solving complex challenges around workload scheduling, resource allocation, cluster management, and infrastructure automation.
This opportunity is ideal for engineers who enjoy building large-scale platforms from the ground up and owning critical infrastructure systems end-to-end.
What You'll DoDesign and build virtualization infrastructure supporting GPU-intensive AI and HPC workloads.
Develop and operate Kubernetes-based orchestration systems for GPU cluster provisioning and workload scheduling.
Build automated provisioning systems that enable GPU capacity to be allocated, scaled, and reclaimed efficiently across multiple tenants.
Design solutions for workload placement, resource management, and cluster lifecycle operations.
Partner closely with hardware, networking, infrastructure, and AI platform teams to ensure orchestration systems align with real-world cluster architectures and constraints.
Improve the reliability, security, scalability, and operational maturity of the orchestration platform.
Build tooling and automation that simplifies infrastructure management and improves developer and customer experiences.
Contribute to architectural decisions, engineering standards, and best practices as the platform evolves.
Strong hands-on experience with Kubernetes and container orchestration in production environments.
Experience designing, building, and operating large-scale infrastructure platforms.
Background with virtualization technologies supporting cloud, HPC, GPU, or distributed computing environments.
Understanding of GPU cluster provisioning, workload scheduling, and resource management.
Experience with Linux-based infrastructure and distributed systems concepts.
Ability to independently own complex systems from design through production operation.
Comfortable working in a fast-moving environment where architecture and processes are being established.
Experience with GPU scheduling technologies such as Slurm, Kubernetes device plugins, NVIDIA GPU Operator, or similar frameworks.
Experience supporting AI infrastructure, machine learning platforms, HPC environments, or GPU cloud providers.
Background building multi-tenant infrastructure platforms for cloud providers or large-scale compute environments.
Experience with infrastructure automation, Infrastructure as Code, and platform engineering…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).