ABOUT EDGE
At EDGE bold ideas are engineered into technologies that protect, improve and save lives.
Headquartered in the United Arab Emirates, EDGE is a leading advanced technology group working at the forefront of defence and emerging technologies. Spanning more than 35 specialised companies and multiple centres of excellence, we are purpose‑built to move fast. Free from heavy legacy processes, we give our people the autonomy, accountability, and agility to bring breakthrough technologies from concept to reality.
Based in Abu Dhabi, a globally connected hub at the crossroads of Europe, Asia, Africa, and the Middle East, EDGE is home to a truly multicultural community where bold ideas thrive, and the future is shaped.
Together, we are shaping the future.
ABOUT THE ROLEWe are seeking a highly skilled HPC Systems Administrator to manage, optimize, and support our high‑performance computing environment used for Computational Fluid Dynamics (CFD) and Finite Element Analysis (FEA) workloads. This role is responsible for the full lifecycle of HPC operations — from infrastructure and scheduler management to user support, performance optimization, and long‑term capacity planning.
The ideal candidate has strong Linux administration experience, deep knowledge of HPC schedulers, and hands‑on familiarity with engineering simulation tools.
- A. Infrastructure Management
- Maintain and administer compute nodes, login nodes, heterogeneous nodes, NAS storage servers, and high‑speed interconnects (Infini Band).
- Manage and maintain HPC‑related databases, ensuring timely backups and audit compliance.
- Monitor hardware health including CPU temperatures, memory errors, and disk failures.
- Ensure high availability, reliability, and minimal downtime across the HPC environment.
- B. Scheduler & Resource Management
- Configure and tune job schedulers (Slurm / PBS / LSF / Grid Engine).
- Implement and maintain fair‑share scheduling, job priority rules, QoS limits, and preemption policies.
- Manage node reservations for large‑scale CFD/FEA workloads.
- Provide best‑practice recommendations for CPU/GPU architecture selection for engineering simulations.
- Prevent resource misuse, including node hogging, queue congestion, and starvation of small jobs.
- C. User Access, Security & Compliance
- Manage user accounts, permissions, and storage quotas.
- Enforce secure SSH access, MFA, and other security controls.
- Ensure compliance with IT governance, data‑security, and audit requirements.
- Apply OS patches, security updates, and vulnerability fixes.
- D. Software Stack Management
- Install, update, and manage licenses for engineering solvers such as ANSYS, Siemens, NASTRAN, and others.
- Maintain version control and apply service pack updates as required.
- Manage environment modules (Lmod, Environment Modules).
- Optimize compilers, MPI libraries, math libraries, and GPU/graphics‑intensive drivers for performance.
- RESTRICTED
- E. Performance Monitoring & Optimization
- Monitor node utilisation, queue wait times, job failures, and overall cluster performance.
- Identify and resolve network congestion, I/O bottlenecks, and memory pressure issues.
- Conduct performance tuning and benchmarking for CFD/FEA workloads.
- Recommend improvements to enhance throughput and efficiency.
- F. Troubleshooting & User Support
- Diagnose and resolve job crashes, memory leaks, solver errors, and environment issues.
- Assist users with HPC batch scripting and job optimization.
- Provide training sessions and best‑practice guidance for HPC users.
- Maintain documentation, onboarding guides, and knowledge‑base resources.
- G. Storage & Data Lifecycle Management
- Enforce storage quotas, purge policies, and data‑retention rules.
- Manage backup, archival systems, and RAID configurations.
- Ensure smooth operation of parallel file systems (Lustre, GPFS, BeeGFS).
- Support efficient data workflows for large CFD/FEA datasets.
- H. Capacity Planning & Future Growth
- Plan for future expansion including increased core counts, GPU adoption, memory‑heavy nodes, and faster interconnects.
- Evaluate new hardware technologies and benchmark workloads before procurement.
- Provide input into long‑term HPC strategy and infrastructure roadmap.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).