×
Register Here to Apply for Jobs or Post Jobs. X

Principal Software Engineer, Rack-Scale System Software — CSP Engagements

Job in Oregon, Dane County, Wisconsin, 53575, USA
Listing for: NVIDIA
Full Time position
Listed on 2026-07-06
Job specializations:
  • Software Development
    Software Engineer, DevOps, Cloud Engineer - Software, Software Architect
Salary/Wage Range or Industry Benchmark: 272000 - 431250 USD Yearly USD 272000.00 431250.00 YEAR
Job Description & How to Apply Below

Overview

We re looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA s cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership.

Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole — fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness.

Responsibilities
  • Drive rack-scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration
  • Drive technical work streams with CSP engineering teams on rack-scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operation
  • Capture and synthesize CSP engineering feedback on rack-scale system software — health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA s architecture decisions
  • Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development
  • Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices — drive documentation, tooling, and test strategy improvements as a result
  • Collaborate with execution teams on left-shift strategy — ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability
  • Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams
Qualifications
  • 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience)
  • Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliability
  • Experience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery)
  • Understanding of error handling and recovery design patterns in distributed systems — fault isolation, retry policies, graceful degradation
  • Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability
  • Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus
  • Customer obsession — genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience
  • Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication — ability to translate complex system software architecture into actionable mentorship for customer engineering teams
Ways To Stand Out From The Crowd
  • Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software
  • Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms)
  • Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices)
  • Familiarity with GPU or…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary