×
Register Here to Apply for Jobs or Post Jobs. X

HPC Systems Engineer

Job in Milpitas, Santa Clara County, California, 95035, USA
Listing for: KLA Corporation
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Systems Engineer, Machine Learning/ ML Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Company Overview

KLA is a global leader in diversified electronics for the semiconductor manufacturing ecosystem. Virtually every electronic device in the world is produced using our technologies. No laptop, smartphone, wearable device, voice-controlled gadget, flexible screen, VR device or smart car would have made it into your hands without us. KLA invents systems and solutions for the manufacturing of wafers and reticles, integrated circuits, packaging, printed circuit boards and flat panel displays.

The innovative ideas and devices that are advancing humanity all begin with inspiration, research and development. KLA focuses more than average on innovation and we invest 15% of sales back into R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work together with the world's leading technology providers to accelerate the delivery of tomorrow's electronic devices. Life here is exciting and our teams thrive on tackling really hard problems.

There is never a dull moment with us.

Group/Division

KLA has always had a close relationship with physics and data. Our optical and electron beam inspection and measurement tools use cutting edge physics models, both for hardware design and as part of their algorithms. AI, including several traditional machine learning techniques and deep learning are routinely used to process this data to meet application requirements. The AI & Modeling Center of Excellence was setup with the mission of advancing KLA's traditional strengths in physics and data and providing implementation solutions for multiple KLA Inspection and Metrology products targeted at the semiconductor manufacturing industry.

The AI & Modeling Center of Excellence is part of the company's Central Engineering organization providing product development expertise in a critical area for a wide variety KLA products. As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex imaging techniques and semiconductor processes.

You will also work with a data scientists and AI infrastructure engineers whose mission is to build and scale machine learning based solutions for our semiconductor customers. We are looking for engineers in a few different fields. If you are passionate about Physics Modeling, High Performance Computing - HPC (including GPU), Machine Learning, Deep Learning, Data Sciences, or cutting-edge Cloud technologies - this is the place for you!

Job Description /Preferred Qualifications

As an AI Infrastructure Engineer, you will research, evaluate, and develop next-generation hardware and software technologies that enable large-scale AI and machine learning workloads across KLA. You will help design and build the infrastructure that powers model training, inference, networking, storage, data protection, and emerging agentic AI systems. This role sits at the intersection of systems engineering, AI platform development, MLOps, and Dev Ops, ensuring that AI teams have reliable, scalable, high-performance infrastructure to develop, train, deploy, and operate AI solutions.

You will also define and develop reference architectures and platform standards that can be embraced across KLA products and engineering organizations, enabling secure, cost-effective, and repeatable AI infrastructure deployments.

Responsibilities

* Design & Build AI

Infrastructure: Architect, deploy, and operate scalable, secure, and cost-effective AI platforms, distributed training environments, and model serving systems.

* Research &

Innovation: Evaluate emerging technologies in AI infrastructure, compute, storage, networking, orchestration, and security, and develop reference designs for enterprise adoption.

* Hardware & Software Orchestration:
Bring up, integrate, and lead AI compute infrastructure while diagnosing hardware, firmware, operating system, and platform issues.

* Data & Storage Management:
Design and maintain high-performance storage, networking, and data pipelines supporting large-scale AI workloads.

* Performance & Reliability:
Develop observability and monitoring solutions, optimize AI workload performance, and ensure infrastructure meets reliability and availability targets.

* Security & Compliance:
Implement security-by-design principles including encryption, identity and access management, secrets management, and AI data security. Contribute to the architecture and implementation capabilities for protecting AI workloads, intellectual property, and sensitive data at KLA.

* Collaboration:

Partner with AI researchers, data scientists, software engineers, IT, and security teams to align platform capabilities with business and product needs.

Tools you will use in the role:

* Technical

Skills:

Linux administration, Kubernetes/Docker, Slurm, Ray, Tensor Flow, PyTorch, GPU/TPU optimization, containerization, orchestration, and programming skills.

* Infrastructure Knowledge:
High…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary