Principal Software Engineer – Rack-Scale System Software
Job in
California, Moniteau County, Missouri, 65018, USA
Listed on 2026-07-20
Listing for:
Jobtailor
Full Time
position Listed on 2026-07-20
Job specializations:
-
Software Development
Software Engineer, DevOps, Software Architect, Cloud Engineer - Software
Job Description & How to Apply Below
Responsibilities
- Drive rack-scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration.
- Drive technical work streams with CSP engineering teams on rack-scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation.
- Capture and synthesize CSP engineering feedback on rack-scale system software — health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA's architecture decisions.
- Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development.
- Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices — drive documentation, tooling, and test strategy improvements as a result.
- Collaborate with execution teams on left-shift strategy — ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability.
- Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams.
- 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering.
- BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience).
- Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability.
- Experience with fabric management software, cluster management, or system-level orchestration frameworks.
- Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery).
- Understanding of error handling and recovery design patterns in distributed systems — fault isolation, retry policies, graceful degradation.
- Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability.
- Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus.
- Customer obsession — genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience.
- Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority.
- Strong communication — ability to translate complex system software architecture into actionable mentorship for customer engineering teams.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×