HPC Architect
Listed on 2026-09-10
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
Company Description
About Abb Vie
Abb Vie's mission is to discover and deliver innovative medicines and solutions that solve serious health issues today and address the medical challenges of tomorrow. We strive to have a remarkable impact on people's lives across several key therapeutic areas including immunology, oncology and neuroscience - and products and services in our Allergan Aesthetics portfolio. For more information about Abb Vie, please visit us at Follow @abbvie on Linked In, Facebook, Instagram, X and You Tube.
Job DescriptionAbb Vie Information Research (IR) Scientific and Cloud Computing team is seeking a highly motivated High Performance Computing (HPC) Architect to provide technical leadership for scientific and R&D computing environments. This role is responsible for architecting and evolving on-premises, cloud, and hybrid HPC platforms, including compute, GPU, networking, storage, workload management, containers, automation, security, and scientific software services. The successful candidate will have strong hands-on experience with HPC infrastructure, AWS, and hybrid cloud environments, Slurm, infrastructure as code, automation, Dev Ops practices, and enterprise IT operations.
The HPC Architect will guide the design, implementation, lifecycle management, and support of secure, scalable, reliable, and cost-effective solutions while partnering with R&D stakeholders and fostering collaboration, innovation, and continuous improvement.
- Serve as the technical architect for Abb Vie IR’s on-premises, cloud, and hybrid HPC based scientific-computing environments. Responsible for architecture, standards, infrastructure (software and hardware) lifecycle, evaluation of POC, capacity planning, and cost management.
- Architect and continuously evolve all aspects of scientific HPC computing platforms, including operational and maintenance processes.
- Design, implement, and support HPC compute and networking infrastructure - compute, networking, and storage.
- Administer and optimize Slurm workload management, including partitions, queues, priorities, fair-share, reservations, resource limits, GPU scheduling, accounting, job troubleshooting, performance tuning, and integration with Posit and other research platforms.
- Establish secure, reproducible container and scientific software environments using Apptainer/Singularity, Docker-compatible workflows, environment modules, compilers, MPI, CUDA, Python, R, and application dependencies.
- Assist with the architecture and support of Posit Workbench, Posit Connect, Posit Package Manager, and related analytics platforms, integrating them with Slurm, Active Directory, storage, GPUs, and scientific software.
- Implement secure identity and access management in accordance with Abb Vie security SOPs and applicable industry standards. Architect, implement, and support vulnerability-management processes and remediation efforts.
- Lead efforts to automate provisioning, configuration, patching, testing, monitoring, remediation, and lifecycle activities.
- Maintain observability metrics, architecture diagrams, inventories, dependency maps, configuration standards, runbooks, support procedures, and end-of-life plans while addressing technical debt, operational risks, and capacity constraints. Lead the development of appropriate reporting and monitoring tools and dashboards.
- Lead design and incident reviews, vendor engagements, proofs of concept, capacity planning, and implementation governance; provide technical direction and mentorship to engineering and partner teams.
- Partner with R&D to create appropriate solutions for research workloads.
- Bachelor’s Degree in Computer Science, IT, or related field with 7 years of experience; OR Master’s Degree with 6 years of experience; OR PhD 2 years' experience.
- Respective years of experience in application program development.
- Leadership experience is required, including the ability to guide, collaborate with and influence technical teams, projects, and work streams both within and outside the organizational structure. Demonstrated ability to balance technical depth with people leadership and stakeholder management.
- Expert knowledge of Linux, HPC compute and GPU architectures, CPU/memory/NUMA/PCIe design, MPI, NVIDIA GPUs, CUDA, MIG, node provisioning, firmware, drivers, and hardware lifecycle management.
- Strong experience with Slurm, including partitions, queues, scheduling policies, resource management, GPU scheduling,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).