Job Description:
About EdgeAt EDGE bold ideas are engineered into technologies that protect, improve and save lives.
Headquartered in the United Arab Emirates, EDGE is a leading advanced technology group working at the forefront of defence and emerging technologies. Spanning more than 35 specialised companies and multiple centres of excellence, we are purpose-built to move fast. Free from heavy legacy processes, we give our people the autonomy, accountability, and agility to bring breakthrough technologies from concept to reality.
Based in Abu Dhabi, a globally connected hub at the crossroads of Europe, Asia, Africa, and the Middle East, EDGE is home to a truly multicultural community where bold ideas thrive, and the future is shaped.
Together, we are shaping the future.
AboutThe Role
We are seeking a highly skilled HPC Systems Administrator to manage, optimise, and support our high-performance computing environment used for Computational Fluid Dynamics (CFD) and Finite Element Analysis (FEA) workloads. This role is responsible for the full lifecycle of HPC operations — from infrastructure and scheduler management to user support, performance optimization, and long-term capacity planning.
The ideal candidate has strong Linux administration experience, deep knowledge of HPC schedulers, and hands‑on familiarity with engineering simulation tools.
Responsibilities- Infrastructure Management
- Maintain and administer compute nodes, login nodes, heterogeneous nodes, NAS storage servers, and high‑speed interconnects (Infini Band).
- Manage and maintain HPC‑related databases, ensuring timely backups and audit compliance.
- Monitor hardware health including CPU temperatures, memory errors, and disk failures.
- Ensure high availability, reliability, and minimal downtime across the HPC environment.
- Scheduler & Resource Management
- Configure and tune job schedulers (Slurm / PBS / LSF / Grid Engine).
- Implement and maintain fair‑share scheduling, job priority rules, QoS limits, and preemption policies.
- Manage node reservations for large‑scale CFD/FEA workloads.
- Provide best‑practice recommendations for CPU/GPU architecture selection for engineering simulations.
- Prevent resource misuse, including node hogging, queue congestion, and starvation of small jobs.
- User Access, Security & Compliance
- Manage user accounts, permissions, and storage quotas.
- Enforce secure SSH access, MFA, and other security controls.
- Ensure compliance with IT governance, data‑security, and audit requirements.
- Apply OS patches, security updates, and vulnerability fixes.
- Software Stack Management
- Install, update, and manage licenses for engineering solvers such as ANSYS, Siemens, NASTRAN, and others.
- Maintain version control and apply service pack updates as required.
- Manage environment modules (Lmod, Environment Modules).
- Optimize compilers, MPI libraries, math libraries, and GPU/graphics‑intensive drivers for performance.
- Performance Monitoring & Optimization
- Monitor node utilization, queue wait times, job failures, and overall cluster performance.
- Identify and resolve network congestion, I/O bottlenecks, and memory pressure issues.
- Conduct performance tuning and benchmarking for CFD/FEA workloads.
- Recommend improvements to enhance throughput and efficiency.
- Troubleshooting & User Support
- Diagnose and resolve job crashes, memory leaks, solver errors, and environment issues.
- Assist users with HPC batch scripting and job optimization.
- Provide training sessions and best‑practice guidance for HPC users.
- Maintain documentation, onboarding guides, and knowledge‑base resources.
- Storage & Data Lifecycle Management
- Enforce storage quotas, purge policies, and data‑retention rules.
- Manage backup, archival systems, and RAID configurations.
- Ensure smooth operation of parallel file systems (Lustre, GPFS, BeeGFS).
- Support efficient data workflows for large CFD/FEA datasets.
- Capacity Planning & Future Growth
- Plan for future expansion including increased core counts, GPU adoption, memory‑heavy nodes, and faster interconnects.
- Evaluate new hardware technologies and benchmark workloads before procurement.
- Provide input into long‑term HPC strategy and infrastructure roadmap.
- Bac…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).