More jobs:
GPU Systems Engineer
Job in
New York, New York County, New York, 10261, USA
Listed on 2026-08-16
Listing for:
Career Techniques
Full Time
position Listed on 2026-08-16
Job specializations:
-
IT/Tech
Systems Engineer, Unix/Linux
Job Description & How to Apply Below
As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes. The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.
Responsibilities:- Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
- Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
- Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
- Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
- Own infrastructure projects end to end, from scope and design through implementation and long-term support.
- Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.
- 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
- Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
- Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
- Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
- Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
- Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
- Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
- Clear communication. You will work daily with researchers, engineers, and vendors.
Comp: 200-300K + Bonus
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×