HPC Systems Administrator
Listed on 2026-10-06
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, Unix/Linux, SRE/Site Reliability
Responsibilities include deploying secure, high-availability AI and HPC systems; automating operations for efficiency; optimizing infrastructure for performance, reliability, and cost-effectiveness; and collaborating with scientists and engineers to facilitate model training and experimentation across GPU, cloud, and on-premises environments.
Requirements:
Bachelor’s in Computer Science or related field; 5+ years supporting large-scale HPC or GPU compute environments; expertise in Linux administration, scripting (Python, Bash), automation tools (Ansible, Kubernetes), and HPC job schedulers (Slurm, Grid Engine). Preferred skills include experience with NVIDIA GPU infrastructure, cloud platforms (AWS, Azure, GCP), and working in regulated or critical environments. The role is based in Silicon Valley with a hybrid work model (3 days onsite, 2 remote).
specifics
Focus on AI/ML workloads, GPU hardware management, distributed computing, and automation for scientific research in drug discovery.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).