More jobs:
Job Description & How to Apply Below
You will automatically engage with top-tier technology as part of a collaborative team overseeing multiple regions.
Your role demands that you set up GPU racks, validate critical networking components, and maintain SLA standards through effective incident response management. The emphasis on hands-on involvement means you will directly impact the efficiency and reliability of our computing services.
Key Responsibilities:
• Build and test new GPU rack configurations
• Optimize Infini Band fabric for peak performance
• Perform capacity planning across regions while forecasting
• Lead incident response efforts to ensure fast resolution
• Create essential runbooks for operational procedures
Requirements:
• Minimum five years of large cluster operation experience
• In-depth knowledge of Infini Band and networking
• Competency in Go or Python programming
• Resilience under pressure during incident management
• Nice to have: experience with provisioning systems
Drive innovation and assurance in compute cluster performance with this role as a Engineer in Site Reliability.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×