IT Administrator — HPC & Data Center Operations
Listed on 2026-08-04
-
IT/Tech
Systems Administrator, Cloud Computing: Infrastructure & Operations, IT Support, Cybersecurity
Job Description:
IT Administrator — HPC & Data Center Operations
Reports To:
Head of IT / Chief Technology Officer
Location:
Abu Dhabi - On-site (Data Center presence required)
Employment Type:
Full-Time
Prepaire Labs is seeking an experienced IT Administrator to manage and
maintain our High-Performance Computing (HPC) environment, including CPU
and GPU data center infrastructure that powers our AI-driven drug discovery and
healthcare research platforms. The successful candidate will be responsible for
end-to-end administration of compute clusters, networking connectivity,
healthcare and scientific software licensing, and day-to-day IT operations,
ensuring high availability, security, and regulatory compliance across all systems
Key Responsibilities 1. HPC & Data Center Administration (CPU/GPU)- Administer, monitor, and maintain HPC clusters comprising CPU and GPU compute nodes (e.g., NVIDIA H100/A100/L40S, AMD EPYC, Intel Xeon platforms).
- Deploy, configure, and manage cluster workload managers and job schedulers (SLURM, PBS, or Kubernetes with GPU orchestration).
Singularity/Apptainer, NVIDIA Container Toolkit), and ML frameworks supporting research workloads.
- Oversee data center operations: rack and stack, power and cooling monitoring (PDU/UPS), cable management, hardware lifecycle, and capacity planning.
- Perform preventive maintenance, firmware/BIOS updates, hardware diagnostics, and coordinate vendor RMA/support cases (NVIDIA, Dell, HPE, enterprise NAS/SAN) and data backup/disaster recovery strategies.
- Monitor cluster health, utilization, and performance using tools such as Grafana, Prometheus, Zabbix, Nagios, or NVIDIA DCGM; optimize resource allocation for research teams.
Design, configure, and maintain LAN/WAN, VLANs, firewalls, VPNs, and wireless infrastructure across office and data center environments.
- Administer high-speed interconnects for HPC workloads (Infini Band, RoCE, 10/25/40/100
GbE).
NVIDIA networking) including switching, routing, and firmware updates.
- Ensure secure remote access, site-to-site connectivity, and redundancy/ failover for critical links.
- Monitor network performance, troubleshoot latency/bandwidth issues, and maintain network documentation, IP address management (IPAM), and topology diagrams.
- Implement network segmentation and access controls appropriate for sensitive healthcare and research data.
- Manage procurement, deployment, renewal, and compliance of software licenses, including scientific, healthcare, and laboratory applications (e.g., LIMS, bioinformatics suites, molecular modeling tools, EHR/EMR integrations where applicable).
- Administer license servers (FlexLM/Flex Net, RLM, etc.) for engineering and scientific software.
- Track license utilization, forecast needs, and optimize licensing costs across teams.
- Ensure all systems handling healthcare or patient-related data comply with applicable regulations and standards (HIPAA, GDPR, ISO 27001, GxP/21 CFR Part 11 as relevant).
- Liaise with healthcare technology vendors, managed service providers, and regulatory/compliance teams.
- Administer Linux (RHEL/Rocky/Ubuntu) and Windows Server environments, including Active Directory / LDAP, DNS, DHCP, and identity/access management (SSO, MFA).
- Implement and enforce security policies: patch management, endpoint protection, vulnerability scanning, intrusion detection, and audit logging.
- Manage virtualization and cloud resources (VMware, Proxmox, Hyper-V; AWS/Azure/GCP hybrid connectivity where required).
- Maintain robust backup, replication, and disaster recovery procedures; conduct periodic DR testing.
- Support incident response, root cause analysis, and change management processes.
- Provide Tier 2/3 support for researchers, scientists, and staff on workstations, peripherals, collaboration tools, and access to HPC resources.
- Onboard/offboard users, manage accounts, permissions, and research project allocations on compute clusters.
- Maintain accurate documentation: SOPs, runbooks, asset inventory, and configuration records.
- Train end users on responsible and efficient use of HPC and IT resources.
- Participate in on-call rotation for critical infrastructure issues.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
- 5+ years of experience in IT systems administration, with 2+ years managing HPC and/or GPU data center environments.
- Strong hands-on expertise in Linux administration (RHEL/CentOS/Rocky/Ubuntu) and shell scripting (Bash, Python).
- Proven experience with GPU computing stacks (NVIDIA drivers, CUDA, DCGM) and job schedulers (SLURM preferred).
- Solid networking knowledge: TCP/IP, VLANs, routing, firewalls, VPNs, and high-speed interconnects (Infini Band/100
GbE a plus). - Experience managing software licensing, license servers, and vendor contracts.
- Familiarity with healthcare data compliance requirements (HIPAA, GDPR) and IT security best…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).