Job Description & How to Apply Below
In this role, you will take primary responsibility for your region's compute clusters and engage in cross-team collaboration. You will be involved in setting up and maintaining GPU racks, ensuring optimal Infini Band fabric performance to meet specifications, and planning for regional capacity needs. This is more than operational oversight; it is a chance to significantly impact the reliability and efficiency of mission-critical systems.
Key Responsibilities:
• Setup and validate GPU racks and configurations
• Ensure Infini Band fabric performance is within required specs
• Run regional capacity planning and demand forecasting
• Manage incident responses, aiming for rapid resolution
• Develop runbooks to enhance team knowledge transfer
Requirements:
• Five-plus years managing large compute clusters
• Strong background in Infini Band and Ethernet networking
• Experience with Go or Python for scripting tasks
• Ability to maintain composure under challenging situations
• Additional familiarity with provisioning systems preferred
Take your skills in cluster management to the next level in this role focused on reliability and performance.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×