More jobs:
Senior Solutions Architect, First Time Deployment Validation - NVIS
Job in
Santa Clara, Santa Clara County, California, 95053, USA
Listed on 2026-09-20
Listing for:
NVIDIA Corporation
Full Time
position Listed on 2026-09-20
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software), Systems Engineer
Job Description & How to Apply Below
You will be embedded in launches from the start, running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters using NCCL and collectives (All Reduce, All To All ) to validate performance and scalability. When workloads or benchmarks fall short, you're the expert who digs in, partners with engineering, and drives resolution. You will operationalize observability and automation to accelerate validation, capture structured evidence across every bring-up milestone, and work directly with internal deployment teams and external customers to ensure AI factories are ready r work directly enables the success of NVIDIA's first external product launches!
** What You Will be Doing:
*** Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
* Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.
* Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
* Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
* Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
* Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks
* Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as All Reduce and All To All .
* Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
* Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.
* Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.
** What We Need to See:
*** Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.
* More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
* Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.
* Solid grasp of collective communication patterns, particularly All Reduce and All To All , and how they are applied in contemporary ML/LLM training.
* Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or Tensor Flow.
* Proficiency with Python and Shell/Bash for scripting, automation, and tooling.
* Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).
* Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.
* Strong communication skills and the ability to work effectively with cross-functional teams.
** Ways to Stand Out From the Crowd:
*** Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.
* Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.
* Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.
* Experience building automation and CI-style pipelines for running and validating benchmarks at scale.
* Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.
NVIDIA is widely considered one of the technology world’s most desirable employers. Some of the world's most forward-thinking and hardworking people are working for us. If you're creative and autonomous, we want to hear from you.#Build
TheAIFactory
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD.You will also be eligible for equity and benefits.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×