×
Register Here to Apply for Jobs or Post Jobs. X

AI HPC Infrastructure Engineer

Job in Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: Analysis Group, Inc.
Full Time position
Listed on 2026-08-02
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 150000 - 170000 USD Yearly USD 150000.00 170000.00 YEAR
Job Description & How to Apply Below

Overview

Analysis Group is one of the largest international economics consulting firms, with more than 1,500 professionals across 15 offices in North America, Europe, and Asia. Since 1981, we have provided expertise in economics, finance, health care analytics, and strategy to top law firms, Fortune Global 500 companies, and government agencies worldwide. Our internal experts, together with our network of affiliated experts from academia, industry, and government, offer our clients exceptional breadth and depth of expertise.

The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux-based clustered computing platform that supports both traditional HPC/analytical workloads and large-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well-utilized, and operations are monitored, documented, and reported — including change management and performance statistics — across both domains.

Essential Job Functions and Responsibilities

  • Maintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.
  • Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi-GPU training strategies (e.g., PyTorch Distributed, Horovod, Deep Speed).
  • Design, deploy, and maintain GPU-accelerated compute infrastructure for large-scale model training and inference.
  • Manage GPU scheduling, multi-tenancy, and utilization across SLURM and/or Kubernetes-based environments.
  • Administer the NVIDIA software stack — drivers, CUDA, cuDNN, NCCL — and coordinate firmware and health monitoring across GPU fleets.
  • Tune and optimize LLM training and inference performance — including batching, quantization, KV-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.
  • Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).
  • Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.
  • Manage access authentication including PAM, LDAP integration, and single sign-on.
  • Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.
  • Manage high-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).
  • Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU-specific issues).
  • Develop and implement backup and recovery programs.
  • Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.
  • Migrate data from heterogeneous environments to Linux, on-prem clusters, or cloud.
  • Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.
  • Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.
  • Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.
  • Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.
  • Ensure continuous network connectivity of all equipment.
  • Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.
  • Participate in a 24x7 on-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.

Qualifications

  • Bachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.
  • A minimum of 5 years of experience as a hands‑on Linux Systems Administrator in a research, HPC, or production setting.
  • An ideal candidate will have 5 to 10 years of substantive relevant experience.
  • Exp…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary