×
Register Here to Apply for Jobs or Post Jobs. X

Research Computing Engineer; AI Infrastructure and HPC)

Remote / Online - Candidates ideally in
State College, Centre County, Pennsylvania, 16801, USA
Listing for: The Pennsylvania State University
Full Time, Remote/Work from Home position
Listed on 2026-09-20
Job specializations:
  • IT/Tech
    Unix/Linux, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software), Systems Engineer
Salary/Wage Range or Industry Benchmark: 91000 - 137000 USD Yearly USD 91000.00 137000.00 YEAR
Job Description & How to Apply Below
Position: Research Computing Engineer (AI Infrastructure and HPC)

POSITION SPECIFICS

Approval of remote and hybrid work is not guaranteed regardless of work location.

For additional information on remote work at Penn State, see Notice to Out of State Applicants.

The Institute for Computational and Data Sciences (ICDS) at Penn State seeks a Research Computing Engineer to join our technical team. This role supports Penn State's research mission by designing, operating, automating, and optimizing the GPU and AI and computing infrastructure used by researchers across the university, along with the high-performance computing systems that support it.

The position will be filled at the Reserach Computing Systems Engineer
- Advanced Professional
level.

Candidates must be U.S. citizens due to specific access requirements associated with this position.

This position is ideal for an engineer who enjoys building reliable, scalable systems for machine learning, data-intensive research, and advanced computing workloads. The successful candidate will work across GPU systems, HPC platforms, storage, networking, automation, and user-facing research workflows to enable cutting‑edge research in AI, simulation, and computational science.

Work Arrangement

This is a full‑time position, which will report to the Research HPC Manager and requires on‑site work at University Park and is not supportive of remote work.

Responsibilities

As part of a collaborative engineering team, you'll contribute across a broad range of responsibilities that include the following:

  • Collaborate with teammates, users, and vendor support to diagnose issues and implement solutions across compute, storage, networking, and software environments.
  • Monitor, maintain, automate, and improve AI and HPC systems and supporting infrastructure.
  • Design, deploy, operate, troubleshoot, and optimize systems using Dev Ops and infrastructure‑as‑code practices.
  • Support GPU‑accelerated computing environments for AI, machine learning, and scientific workloads.
  • Partner with researchers and ICDS staff to understand workload requirements and develop practical engineering solutions for system configuration, performance, and research workflows.
  • Support security, logging, documentation, and compliance process for the systems we operate, including environments subject to federal research security.
  • Contribute to planning, requirements gathering, process improvement, and operational readiness for new services and infrastructure.
  • Provide timely updates to system documentation and respond to user questions with clear, actionable guidance.
  • Evaluate and improve tools, platforms, and workflows that support AI model development, training, inference, and data movement at scale.
Required qualifications and skills include the following
  • Administration of multi‑GPU nodes at scale, including driver and firmware lifecycle management, NVLink/NVSwitch topology validation, GPU health monitoring and tuning for multi‑GPU or multi‑node GPU workloads.
  • Ability to work effectively in a Linux environment, including command‑line tools, file editing, POSIX permissions, and system configuration.
  • Strong scripting ability in Bash and Python.
  • Strong problem‑solving skills and the ability to debug complex systems. Ability to work collaboratively as part of a technical team. Experience using AI tools or AI agents to improve programming, debugging, development, or prototyping workflows.
  • Clear written and verbal communication skills.
Preferred qualifications

Experience with one or more of the following is helpful but not required:

AI and HPC Workloads
  • Experience supporting AI/ML infrastructure, including environments used for model training, inference, experiment workflows, and large‑scale data processing.
  • Experience with job schedulers such as Slurm, PBS, HTCondor, or LSF.Software…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary