HPC Architect / Lead
Listed on 2026-08-09
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Owns the technical strategy and design of the client’s hybrid HPC platform — responsible for what we build and why. The Architect defines the platform roadmap, leads technology evaluation, designs the hybrid cloud architecture, and ensures the platform meets the performance, security, and cost requirements of seismic processing and AI workloads. This role operates at the intersection of deep technical expertise and enterprise strategy, engaging with geoscience stakeholders, IT leadership, vendors, and finance to align platform capability with business demand.
Key Responsibilities:
1. Design the hybrid HPC platform architecture: on-premises NVIDIA H200 GPU compute integrated with scalable cloud infrastructure (AWS/GCP, to be selected).
2. Define and own the multi-year HPC platform roadmap, aligned to corporate strategy and geoscience business requirements.
3. Own HPC capability management, including current skills assessment and future capability planning.
4. Forecast medium- and long-term resource demand for the discipline in collaboration with other functions as needed.
5. Ensure appropriate deployment of resources across projects to balance delivery, development, and well-being.
6. Identify skills gaps and develop actions to close them through hiring, development, or external support
7. Sponsor and govern performance, development, and well-being outcomes across the HPC discipline
8. Lead and own annual HR people processes for the discipline, including but not limited to:
Succession planning, Talent and workforce planning, Annual salary and reward processes, Approval of new roles and hiring decisions, Oversight of HPC-related expenses and people costs etc.
9. Provide clear, consistent line management and people leadership for team members.
10. Lead performance management, goal setting, feedback, and development planning.
11. Coach and develop as needed, building leadership capability within the team.
12. Foster a culture of psychological safety, accountability, learning, and inclusion.
13. Manage succession planning and retention of critical skills.
14. Role model and uphold the leadership behaviours outlined in the Behavioural Framework.
15. Sponsor and support technical learning, mentoring, and knowledge sharing across the team.
16. Encourage innovation, continuous improvement, and curiosity within the team.
17. Create opportunities for cross-disciplinary collaboration and learning.
18. Lead technology evaluation and vendor selection for storage (VAST, WEKA, DDN), cloud services, networking, and GPU compute evolution.
19. Design hybrid workload orchestration: policies determining where jobs execute based on cost, data locality, capacity, and performance requirements.
20. Architect cloud networking (VPN, dedicated interconnect, peering) and hybrid data movement pipelines between on-prem and cloud.
21. Define the AI/ML infrastructure strategy: platform capabilities required to support distributed training, inference, and MLOps at scale.
22. Establish Fin Ops practice for HPC: cloud cost visibility, chargeback/showback models, reserved capacity planning, and cost governance.
23. Design security and compliance architecture for the hybrid platform aligned to ISO 27001:2022, including multi tenancy, data sovereignty, access control, and firmware supply chain integrity.
24. Lead GPU-aware scheduling strategy: partition design, resource allocation policies, and capacity planning for mixed seismic and AI workloads.
25. Engage geoscience, seismic operations, IT, and finance stakeholders to translate business demand into platform capabilities and investment cases.
26. Manage vendor relationships with NVIDIA, storage vendors, and cloud providers at a technical and commercial level.
27. Own architecture governance: design review, technology standards, technical debt management, and exception handling.
28. Guide and technically direct the Principal HPC Engineer, ensuring implementation aligns with architectural intent.
29. Contribute to budget planning, procurement strategy, and investment cases for HPC capital and operational expenditure. 30. Identify and manage technical risks to platform delivery, performance, and continuity;…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: