More jobs:
Senior Manager, Kubernetes Runtime Engineering
Job in
Santa Clara, Santa Clara County, California, 95052, USA
Listed on 2026-10-01
Listing for:
Nvidia
Full Time
position Listed on 2026-10-01
Job specializations:
-
IT/Tech
Systems Engineer
Job Description & How to Apply Below
You will work across networking, storage, GPU resource management, and cluster security to deliver a production-grade, multi-tenant Kubernetes platform. Your team's decisions directly shape the runtime foundation that internal and external customers depend on.
What You'll Be Doing:
Be responsible for the build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all supported topologies
Manage a team of engineers coordinating the entire container runtime stack: AICR, GPU management operator, DCGM, and related node-level components
Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA , and GPU resource partitioning (MIG, MPS, time-slicing)
Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries
Partner with NKE platform, infrastructure, and cybersecurity teams to integrate new capabilities and resolve cross-cutting runtime concerns
Build and maintain tooling for AICR lifecycle management — provisioning, upgrades, configuration drift detection, and remediation
Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership
Contribute to open source communities anywhere NKE has upstream dependencies or influence
What We Need to See:
BS/MS degree in Computer Science or related field (or equivalent experience)12+ overall years of relevant experience designing and delivering large-scale distributed software systems, including 5+ years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software.
Experience leading a group of engineers with varying specializations and seniority levels — bridging runtime, networking, and security fields is a core part of this role Kubernetes internals knowledge — not just usage; you understand how the scheduler, kubelet, API server, and admission controllers interact
Cluster lifecycle management experience — Cluster API, kubeadm, or equivalent; experience leading fleet-scale cluster provisioning and upgrades
Security and compliance posture — CIS Kubernetes Benchmark, pod security admission, image signing, supply chain integrity
Proven ability to design and implement maintainable APIs for consumers
Familiarity with Identity and Access Management approaches
Excel in managing up, down, and across organizations
Demonstrated ability to reach cross-organization consensus without all the details
Ways to Stand Out from the crowd:
Prior experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling
Experience running Kubernetes at hyperscale with GPU node pools
Track record of upstream open source contributions in the Kubernetes or any open source runtime ecosystem
Experienced, persuasive, and effective interpersonal skills — written, verbal, and in front of engineering leadership
Demonstrated skills in coaching, analysis, problem solving, and short/long-term technical planningNVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×