More jobs:
Senior System Engineer – GPU Platforms
Job in
San Jose, Santa Clara County, California, 95199, USA
Listed on 2026-09-12
Listing for:
Jobtailor
Full Time
position Listed on 2026-09-12
Job specializations:
-
IT/Tech
Unix/Linux, Systems Engineer, Hardware Engineer
Job Description & How to Apply Below
- Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms
- Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes
- Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components
- Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation
- Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities
- Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution
- Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering teams, and external technology partners
- Develop technical documentation, troubleshooting guides, and best practices
- Deliver technical presentations, training sessions, and internal knowledge-sharing activities
- Serve as a technical resource and mentor for other engineers when appropriate
- Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience
- 5–15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field
- Strong knowledge of enterprise server hardware and system architecture
- Hands-on experience with Linux server environments
- Experience installing, configuring, validating, and troubleshooting server hardware and software
- Strong system-level troubleshooting and root-cause-analysis skills
- Working knowledge of PCIe architectures and high-performance I/O
- Experience with GPU computing, accelerators, or comparable high-performance computing technologies
- Ability to independently manage complex technical assignments and drive issues toward resolution
- Strong written and verbal communication skills
- Ability to work effectively with cross-functional and geographically distributed engineering teams
- Comfortable participating in customer-facing technical discussions
- Preferred:
Hands-on experience with NVIDIA data center or professional GPU platforms - Preferred:
Experience with CUDA and NVIDIA GPU software environments - Preferred:
Experience with NVIDIA NVQUAL or similar platform qualification processes - Preferred:
Experience with 4-GPU or 8-GPU server platforms - Preferred:
Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures - Preferred:
Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication - Preferred:
Experience with GPUDirect RDMA, Infini Band, RoCE, or high-speed Ethernet - Preferred:
Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools - Preferred:
Experience with Docker, containers, Kubernetes, or related orchestration technologies - Preferred:
Experience supporting AI, machine learning, HPC, or accelerated computing environments - Preferred:
Experience with customer POCs, technical evaluations, or engineering escalations - Preferred:
Experience delivering technical training or knowledge-sharing sessions - Bash, Python, or other scripting experience is a plus
Demonstrates expertise in GPU server platform support, including installation, configuration, and troubleshooting of Linux environments and GPU technologies. Proficient in system-level diagnostics, technical documentation, and cross-functional collaboration to drive complex technical assignments to resolution.
Highest-signal resume keywords- GPU Computing
- Linux Server Environments
- System-Level Troubleshooting
- NVIDIA NVQUAL
- Technical Documentation
- GPU Drivers
- CUDA
- Server Hardware Configuration
- Root-Cause Analysis
- PCIe Architectures
- High-Performance Computing
- NVIDIA Data Center Platforms
- Docker
- Python Scripting
- Bash Scripting
- Strong Communication Skills
- Cross-Functional Collaboration
- Customer-Facing Technical Discussions
- Mentoring
- Systems Engineering
- Server Engineering
- Platform Engineering
- Technical Enablement
- AI Infrastructure
- HPC
- NCCL
- NVML
- DCGM
- Fabric Manager
- Infini Band
- RoCE
- High-Speed Ethernet
- Kubernetes
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×