Senior HPC Engineer
Listed on 2026-08-23
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure, Network Engineer
Core
42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core
42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.
We are seeking a highly skilled Senior HPC Engineer to support the design, implementation, deployment, and ongoing operations of high-performance computing infrastructure. The role will be responsible for ensuring the availability, reliability, and performance of complex HPC environments spanning compute, storage, networking, Infini Band, GPU resources, job scheduling, and supporting platforms.
The ideal candidate will bring strong hands‑on experience across Linux systems administration, HPC cluster deployment, network engineering, storage, workload management, and automation. You will work closely with internal engineering teams, customers, and technology vendors to troubleshoot complex issues, optimize infrastructure, and maintain highly available HPC environments supporting demanding workloads, including AI/ML and specialized industry applications.
Your key responsibilities- Support the design, implementation, operation, and maintenance of Core
42’s HPC infrastructure, including compute, storage, networking, Infini Band, and associated management platforms. - Install, configure, deploy, and maintain HPC clusters, including compute nodes, storage nodes, interconnects, and supporting infrastructure.
- Configure, manage, and troubleshoot Infini Band and high-speed Ethernet networks, including NVIDIA Mellanox switches, subnet managers such as OpenSM, and routing configurations.
- Implement and maintain job scheduling and workload management platforms, with strong experience in LSF preferred and exposure to Slurm and/or PBS.
- Administer and support Red Hat Enterprise Linux and Windows Server environments within complex HPC and data center infrastructures.
- Support HPC storage environments, including IBM Spectrum Scale (GPFS), DDN Grid Scaler, Net App SAN/NAS, and backup solutions such as IBM Spectrum Protect (TSM).
- Utilize HPC administration and provisioning toolkits, including xCAT, PXE, Kickstart, and other automated deployment technologies.
- Support and optimize GPU-enabled infrastructure, including NVIDIA H100 or newer architectures, and tune workloads for CPU and GPU resources.
- Monitor system, network, storage, and infrastructure health using tools such as Zabbix, Grafana, and other monitoring and observability platforms.
- Perform advanced troubleshooting and root cause analysis across complex HPC, network, storage, and operating system environments.
- Support virtualization platforms including VMware ESXi/vSAN, Citrix, VxRail, and KVM where required.
- Administer and support enterprise server platforms, including HPE Pro Liant, Dell Power Edge, and other compute infrastructure.
- Develop automation and operational tooling using Shell scripting and Python to improve efficiency, reliability, and repeatability.
- Create and maintain technical documentation, architecture diagrams, and operational procedures using tools such as Visio and Draw.io.
- Apply HPC security best practices, including encryption, firewalls, access controls, network security, and compliance requirements.
- Collaborate effectively with customers, vendors, and internal multidisciplinary teams to resolve technical issues and support infrastructure improvements.
- Support specialized HPC workloads and, where applicable, industry applications such as reservoir engineering simulators and petroleum software tools including tNavigator, Nexus-VIP, Eclipse, Intersect, IMPOWER, and PUMA.
- Participate in continuous improvement initiatives to enhance the performance, scalability, availability, and reliability of HPC environments.
What we’re looking for (a) Required skills / qualifications
- Bachelor’s or Master’s degree in Computer Science,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).