More jobs:
Job Description & How to Apply Below
Drive innovation in AI Infrastructure at NVIDIA as a Senior HPC Infrastructure Engineer. Contribute your expertise in high-performance computing, GPU technology, and networking in a dynamic environment.
NVIDIA is seeking a highly skilled Senior HPC Infrastructure Engineer with over 10 years of experience in software engineering for large-scale systems.
Your role will involve managing end-to-end automation of data center operations, enhancing reliability, and ensuring seamless integration across platforms. You’ll collaborate with multidisciplinary teams to tackle complex challenges and drive advancements in machine learning infrastructure.
Key Responsibilities:
• Automate GPU provisioning, configuration, and lifecycle management
• Implement monitoring for reliability and scalability of GPU assets
• Manage NVLINK topology across GPU clusters
• Develop automated test infrastructure for distributed systems
• Collaborate with engineering teams on software integration
Requirements:
• 10+ years in large-scale software engineering
• BS in Computer Science, Engineering, or related fields
• Proficient in Go or Python, data structures, and algorithms
• Strong Linux system administration skills
• Familiar with cluster management systems like Kubernetes
Use your expertise to shape the future of GPU infrastructure at NVIDIA.
#J-18808-Ljbffr
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×