Software Engineer; Infrastructure
Listed on 2026-08-24
-
Software Development
DevOps, Cloud Engineer - Software, Unix/Linux
Location: Northern
Company Thunder Compute is building the VMware for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic. Deployed GPU fleets are currently only 5-20% utilized. Leading solutions for under utilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads;
hence, it must sit at the systems layer. We are a team of systems researchers product ionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer. Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through , to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center.
This enables something like "Ceph for GPUs": GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.
Your work will focus on building the cloud infrastructure surrounding our GPU virtualization layer. This includes the Go backbone of our cloud platform, Kubernetes-based orchestration, production reliability, networking, storage, billing infrastructure, and the systems used to deploy and operate GPU capacity will take ownership of complex infrastructure from early design through production deployment. Example projects may include:
- Building control-plane services for provisioning and managing virtual GPU instances
- Designing reliable systems for GPU allocation, scheduling, and lifecycle management
- Improving our unconventional Kubernetes deployment, which acts as a form of hypervisor for customer workloads
- Building infrastructure for networking, storage, authentication, billing, and usage metering
- Automating the deployment and operation of GPU hosts across cloud providers and customer data centers
- Debugging failures across customer workloads, Kubernetes, our control plane, the network, and physical GPU infrastructure
- Designing systems for failure recovery, capacity management, observability, and incident response
- Improving the security, reliability, and operational simplicity of the platform as it scales
- Working directly with customers to diagnose problems and deploy Thunder Compute in new environments
You will spend your days bouncing between the weeds of complicated production infrastructure that is live and used by customers. One week, you may be debugging a networking failure across a Kubernetes cluster; the next, you may be redesigning the provisioning system to make deployments faster and more reliable. This work is not easy. It blends the hardest parts of cloud infrastructure, distributed systems, and production engineering.
We look for exceptional engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable infrastructure.
- Exceptional Go ability, including concurrency, distributed systems design, API design, and production service development
- Deep understanding of Kubernetes, containers, Linux, networking, storage, or cloud infrastructure
- Experience building and operating critical production systems
- Strong systems debugging and operational ability
- Ability to reason through unfamiliar systems across multiple layers of the stack
- Working knowledge of Python; familiarity with Type Script or Next.js is helpful
- Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines
- Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight
- Strong ownership over correctness, reliability, performance, and operational outcomes
- Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner
- Willingness to work directly with customers and investigate difficult production failures
- Strong communication skills and the ability to coordinate across engineering, customers, and external infrastructure providers
- Experience with Kubernetes internals, container runtimes, cloud networking, distributed storage, infrastructure security, or large-scale control planes
- Experience building high-stakes production infrastructure at a trading firm such as Citadel Securities or Jane Street; a cloud provider such as AWS, Core Weave, or Lambda; an AI infrastructure company; or a similarly demanding engineering environment
- Strong computer science fundamentals demonstrated through academic work, distributed systems research, open-source contributions, or exceptional professional experience
- Experience designing and operating infrastructure across multiple cloud providers or on-premise environments
- Experience taking a new infrastructure system…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).