Job Description & How to Apply Below
Help shape NVIDIA DGX Cloud's automation and tooling as a Senior Software Engineer. Join a production engineering team focused on Kubernetes-based infrastructure and large-scale GPU cluster operations.
You'll be instrumental in developing reliable and scalable systems that enhance the operational life cycle of GPU clusters. Collaborate with multidisciplinary teams to improve workflows and reduce manual interventions using APIs and Git Ops practices. This role involves a strong focus on Kubernetes, automation, and ensuring production readiness across varied environments.
Key Responsibilities:
• Build automation for GPU clusters across cloud and on-prem environments
• Develop tools for lifecycle management, validation, and monitoring
• Enhance workflows for cluster bringup and operations
• Automate tasks to minimize manual production interactions
• Participate in on-call duties and incident response
Requirements:
• 8+ years in production infrastructure engineering
• Proficiency in Python, Go, or similar languages
• Hands-on experience with Linux, Kubernetes, and cloud infrastructure
• Ability to troubleshoot distributed systems
• BS/MS in Computer Science or related field
Leverage your expertise in automation, Kubernetes, and GPU operations to advance NVIDIA’s leading cloud offerings.
#J-18808-Ljbffr
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×