Systems Administrator Lead; Second Shift
Listed on 2026-09-04
-
IT/Tech
Systems Administrator, Unix/Linux, Systems Engineer, IT Support
Join the Future of Digital Infrastructure
Are you a passionate Systems Administrator Lead (Second Shift) looking to make a meaningful impact? We're building the next generation of digital infrastructure powering hyperscalers, AI innovation, and high-performance computing across North America.
Role SummaryAs a Systems Administrator Lead covering the Second Shift (4PM– 1AM EST), you will support, maintain, troubleshoot, and optimize large-scale HPC and AI infrastructure across our data center and cloud environments.
This role is primarily focused on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management. This position provides critical overnight coverage, ensuring infrastructure issues are resolved quickly outside standard business hours.
You will serve as a senior technical escalation point for complex infrastructure incidents affecting GPU clusters, compute nodes, networking, storage, and supporting management services, working closely with Network Engineering, Data Center Operations, Deployment Engineering, vendors, and customer technical teams to restore service and improve the reliability of production HPC environments.
How We Work At 5COur core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great.
What You Will Do HPC Infrastructure Operations- Administer, maintain, and troubleshoot Linux-based HPC and AI compute environments, including large-scale GPU clusters built on NVIDIA HGX, DGX, or equivalent accelerated computing platforms.
- Diagnose hardware, OS, driver, firmware, networking, and storage-related failures, and perform root-cause analysis to develop corrective actions for recurring issues.
- Serve as a senior escalation point for complex incidents affecting production compute infrastructure, troubleshooting degraded or unstable compute/GPU nodes to restore service.
- Support infrastructure across heterogeneous environments - bare-metal, virtualized, containerized, and cloud-hosted - and maintain operational documentation, runbooks, and shift-handoff notes.
- Install, configure, validate, and troubleshoot NVIDIA GPU drivers, CUDA components, firmware, and supporting system software using tools such as Nvidia Smi, DCGM, and NVIDIA Fabric Manager.
- Investigate GPU Xid errors, NVLink/NVSwitch faults, PCIe errors, ECC events, GPU resets, and thermal or power-related issues.
- Diagnose communication and performance issues involving NCCL, GPUDirect RDMA, PCIe topology, NUMA placement, and multi-GPU systems, validating GPU-to-GPU, GPU-to-network, and GPU-to-storage communication after repairs.
- Coordinate replacement and RMA activities for GPUs, GPU trays, baseboards, NVSwitch components, system boards, and NICs.
- Administer enterprise Linux distributions across the production fleet, troubleshooting boot failures, kernel issues, systemd services, file systems, and OS performance.
- Analyze system and kernel logs (journalctl, dmesg, lspci, dmidecode, ipmitool, ethtool, ss, sar, perf) and manage kernel modules, device drivers, DKMS packages, and OS updates.
- Support Linux networking, storage mounts, authentication, and security controls, and investigate OS crashes, kernel panics, and out-of-memory conditions.
- Support enterprise server platforms from vendors such as Dell, HPE, ASUS, Supermicro, and NVIDIA, troubleshooting processors, memory, PCIe devices, NICs, storage controllers, and other hardware components.
- Manage BIOS, BMC, CPLD, NIC, GPU, switch, and drive firmware, and perform remote console troubleshooting, power cycling, and log bundle generation.
- Partner with on-site Data Center technicians on component replacement, cabling validation, break-fix activity, and post-repair testing.
- Troubleshoot Ethernet, Infini Band, and RoCE connectivity within HPC and AI environments, including link state, VLAN/bonding issues, MTU mismatches, packet loss, and RDMA/GPUDirect RDMA…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).