×
Register Here to Apply for Jobs or Post Jobs. X

HPC Systems Engineer

Job in Ann Arbor, Washtenaw County, Michigan, 48113, USA
Listing for: Dormont Manufacturing Co
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    Systems Engineer
Salary/Wage Range or Industry Benchmark: 105900 - 180000 USD Yearly USD 105900.00 180000.00 YEAR
Job Description & How to Apply Below

Company Overview

KLA is a global leader in diversified electronics for the semiconductor manufacturing ecosystem. Virtually every electronic device in the world is produced using our technologies. No laptop, smartphone, wearable device, voice‑controlled gadget, flexible screen, VR device or smart car would have made it into your hands without us. KLA invents systems and solutions for the manufacturing of wafers and reticles, integrated circuits, packaging, printed circuit boards and flat panel displays.

The innovative ideas and devices that are advancing humanity all begin with inspiration, research and development. KLA focuses more than average on innovation and we invest 15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem‑solvers work together with the world’s leading technology providers to accelerate the delivery of tomorrow’s electronic devices. Life here is exciting and our teams thrive on tackling really hard problems.

There is never a dull moment with us.

Group/Division

With over 40 years of semiconductor process control experience, chipmakers around the globe rely on KLA to ensure that their fabs ramp next‑generation devices to volume production quickly and cost‑effectively. Enabling the movement towards advanced chip design, KLA’s Global Products Group (GPG), which is responsible for creating all of KLA’s metrology and inspection products, is looking for the best and the brightest research scientist, software engineers, application development engineers, and senior product technology process engineers.

Central Engineering is KLA’s largest engineering organization comprised of 9 Centers‑of‑Excellence (CoE) in various disciplines applied across all product groups in the company. These CoE include Handling & Automation, Precision Motion Control, Sensors & Image Acquisition, Platform Design, and Packaging Engineering, among others. Talent includes over 500 engineers across global centers in Israel, China, India, and the US. Each CoE contributes not just talent and deliverables per discipline toward product programs, but also subject matter expertise, best practices, roadmaps, specialized facilities, apparatus, models, and analytics.

These differentiate KLA not only in WHAT we do, but also in HOW we do it.

Job Description /

Preferred Qualifications

We’re looking for a HPC Systems Engineer to help power the compute infrastructure behind our R&D innovation! In this role, you’ll support and evolve a high‑performance Linux cluster used for physics modeling, simulation, algorithm development, and machine‑learning workloads—enabling hundreds of engineers to do their best work every day. You’ll play a key role in driving the reliability, performance, and scalability of a shared, mission‑critical HPC environment, partnering closely with infrastructure, Dev Ops, and application teams to keep the platform fast, resilient, and ready for the most demanding computational challenges!

Key Responsibilities HPC Platform Operations
  • Operate and maintain a large‑scale Linux based HPC cluster used for internal R&D workloads
  • Manage compute nodes, login nodes, and supporting infrastructure in a multi‑tenant environment
  • Monitor cluster health, performance, and capacity; respond to incidents and degradations
Scheduler & Workload Management
  • Configure, tune, and support HPC job schedulers (e.g., SLURM, LSF, PBS, or equivalent)
  • Assist users with job submission issues, resource requests, and queue optimization
  • Help optimize scheduler policies to balance throughput, fairness, and utilization
Linux Systems Engineering
  • Install, configure, and maintain Linux operating systems across compute and service nodes
  • Manage OS updates, kernel changes, drivers (including GPU drivers where applicable), and system hardening
  • Troubleshoot complex Linux performance, networking, storage, and process level issues
Performance & Scaling
  • Support high throughput and parallel workloads across CPU and GPU resources
  • Diagnose performance bottlenecks across compute, storage, network, and scheduler layers
  • Assist with scaling activities such as node expansions, re provisioning, and hardware refreshes
Automation &…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary