Global Factory Systems Engineering Manager - Diagnostics
Listed on 2026-07-14
-
Software Development
Software Engineer, Software Project Mgr/ Lead, DevOps, AI Engineer (Applied/Software)
NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.”We'relooking
to grow our company and establish teams with the most thoughtful people in the world.
We are the Datacenter System Software team, and we are looking for a highly motivated, creative Engineering Manager to drive Factory System Software and Diagnostics Integration end to end. You will build and lead a global engineering team delivering embedded code, application programs, and diagnostic updates. These updates support factories building NVIDIA's GPU- and DPU-based products. This includes tightly coupled rack-scale systems such as GB200/GB300 NVL
72 and next-generation platforms. The work covers concurrent NPI ramps and sustaining production. You will partner with system architects, firmware developers, SWQA, product engineering, compliance and security teams, program and product management, and ODM/CM manufacturing partners to ensure the highest-quality releases land on factory floors — and that no bug is discovered there first.
Join us at the forefront of technological advancement.
What you’ll be doing:
Build, lead, mentor, and grow a global factory engineering team spanning the US and Taiwan — operating a follow-the-sun coverage model with on-site presence at ODM/CM partner factories. Own hiring, career development, calibration, and succession planning.
Define Factory readiness scope and workflows for rack scale products coordinating multi-functionally with product management, technicalarchitectsand program management. Deliver those workflows through the validation matrix, ensuring delivered firmware and software is of the highest quality. Solutions must scale and be resilient.
Own technical leadership for how firmware, software, and diagnostics releases reach factories building rack-scale systems. These systems include tightly coupledcomputeand switch trays. Build the end-to-end infrastructure and workflows that ensure every release arrives with efficient quality.
Left-shift release quality: partner with all matrixed organizations — developers, SWQA, and product engineering — in a fast-moving environment with end-to-end CI/CD so that no bug is first found at a factory site. Enforce well-placed quality gates at every product landmark, publish and track indicators at a regular cadence, and report release progress to collaborators and executives.
Own the factory escalation path: triage SLAs, 24×7 coverage, failure root-causeand deflection, andbonepileburn-down — minimizing line-down time through NPI ramps and mass production.
Shape the team's roadmap and drive innovation with a strong focus on automation and AI-assisted validation and triage — automating station readiness, firmware-update flows, and log triage so senior engineering time shifts from setup to analysis.
Continuously analyze factory processes, systems, and workflows toidentifyimprovement and optimization opportunities; remove bottlenecks, document and publish standard operating procedures (SOPs), and ensure the team performs in the most efficient and transparent way against measurable targets.
What we need to see:
10+ overall years in the software industry with specialization in system software and/or firmware development.
3+ years of engineering management or technical leadership experience, including building and leading geographically distributed teams.
BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
Proven track recordof shipping scalable server products through factory ramps — from NPI bring-up to mass production — collaborating with hardware, firmware, manufacturing, diagnostics, and QA teams.
Experience working with ODM/OEM partners to deliver quality servers and solutions for large-scale data centers.
A self-starter who loves finding creative solutions to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).