Senior Network Engineer – GPU Cluster Networking
Listed on 2026-08-30
-
IT/Tech
Systems Engineer, Network Engineer
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.
Whether you’redesigning next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger— technology that moves the world forward.
Join us and, together, we’ll advance your career.
We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.
This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.
The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.
The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.
You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.
THE PERSON:You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.
You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.
You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.
KEY RESPONSIBILITIES:- Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
- Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
- Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
- Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
- Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
- Perform network topology modeling, over subscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
- Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
- Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
- Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
- Lead production incident response,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).