×
Register Here to Apply for Jobs or Post Jobs. X

Senior Systems Software Engineer - Fleet Debuggability

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: NVIDIA
Full Time position
Listed on 2026-07-24
Job specializations:
  • Software Development
    DevOps, Python, Software Engineer, Unix/Linux
Salary/Wage Range or Industry Benchmark: 184000 - 356500 USD Yearly USD 184000.00 356500.00 YEAR
Job Description & How to Apply Below

We are the Datacenter System Software team, and we are looking for a highly motivated, creative Senior Engineer to drive Fleet Scale Debuggability end to end. You will design, architect, and build infrastructure, tooling, analytics on how to collect multi-rack scale logs. The solution should normalize, correlate, and reason over logs spanning multiple components, trays, or racks including NVIDIA's GPUs, CPUs, Network products.

The logs shall be fetched inband or out of band and should help triage fleet level issues seen by our customers. Your work directly shortens the path from a raw, noisy log stream to an actionable root cause. Join us at the forefront of technological advancement.

What you will be doing:
  • Architect, Design, build, fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time-align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in-band and out-of-band channels. Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, so triage is repeatable rather than tribal knowledge.
  • Develop debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults. Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
  • Partner with all matrixed organizations — developers, SWQA, and product engineering — in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
  • Steward the project's open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
  • Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.
What we need to see:
  • 10+ years in the software industry with specialization in system software and/or firmware development.
  • BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
  • Proven track record of shipping scalable server products or fleet-wide experience.
  • A self-starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills — including executive-level reporting — strong work ethic, and dedication to teamwork.
  • Flexibility to work and communicate effectively across teams, partners, and time zones.
  • Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira. Strong, demonstrable skills in Python or RUST.
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
  • Hands‑on experience with out‑of‑band management and platform interfaces — BMC, Redfish, IPMI, SEL — and an understanding of in‑band vs. out‑of‑band trade‑offs.
  • Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine‑readable events.
Ways to stand out from the crowd:
  • Experience leading debuggability solutions on sophisticated rack‑scale compute architectures like GB200/GB300 NVL
    72. Familiarity with log and telemetry analytics stacks (e.g., Open Search/ELK, Loki, Prometheus, Grafana, Pager Duty) and time‑series databases.
  • Hands‑on experience with x86/ARM system architecture and coding (C/C++, Python). Experience with SCM (Git, Perforce) and project management tools (Jira).
  • Track record of integrating AI/LLM tooling into engineering workflows — for triage, validation, log analysis, or test generation. Experience standing up follow‑the‑sun support organizations with measurable response SLAs
  • Experience contributing to…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary