×
Register Here to Apply for Jobs or Post Jobs. X

Systems Administrator Lead; Second Shift

Job in Springfield, Clark County, Ohio, 45502, USA
Listing for: 5C
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    Systems Administrator, Unix/Linux, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 115000 - 145000 USD Yearly USD 115000.00 145000.00 YEAR
Job Description & How to Apply Below
Position: Systems Administrator Lead (Second Shift)

Join the Future of Digital Infrastructure

Are you a passionate Systems Administrator Lead (Second Shift) looking to make a meaningful impact? We're building the next generation of digital infrastructure powering hyperscalers, AI innovation, and high-performance computing across North America.

Role Summary

As a Systems Administrator Lead covering the Second Shift (4PM– 1AM EST), you will support, maintain, troubleshoot, and optimize large-scale HPC and AI infrastructure across our data center and cloud environments.

This role is primarily focused on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management. This position provides critical overnight coverage, ensuring infrastructure issues are resolved quickly outside standard business hours.

You will serve as a senior technical escalation point for complex infrastructure incidents affecting GPU clusters, compute nodes, networking, storage, and supporting management services, working closely with Network Engineering, Data Center Operations, Deployment Engineering, vendors, and customer technical teams to restore service and improve the reliability of production HPC environments.

How We Work At 5C

Our core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great.

What You Will Do HPC Infrastructure Operations
  • Administer, maintain, and troubleshoot Linux-based HPC and AI compute environments, including large-scale GPU clusters built on NVIDIA HGX, DGX, or equivalent accelerated computing platforms.
  • Diagnose hardware, OS, driver, firmware, networking, and storage-related failures, and perform root-cause analysis to develop corrective actions for recurring issues.
  • Serve as a senior escalation point for complex incidents affecting production compute infrastructure, troubleshooting degraded or unstable compute/GPU nodes to restore service.
  • Support infrastructure across heterogeneous environments - bare-metal, virtualized, containerized, and cloud-hosted - and maintain operational documentation, runbooks, and shift-handoff notes.
GPU and Accelerated Computing Systems
  • Install, configure, validate, and troubleshoot NVIDIA GPU drivers, CUDA components, firmware, and supporting system software using tools such as Nvidia Smi, DCGM, and NVIDIA Fabric Manager.
  • Investigate GPU Xid errors, NVLink/NVSwitch faults, PCIe errors, ECC events, GPU resets, and thermal or power-related issues.
  • Diagnose communication and performance issues involving NCCL, GPUDirect RDMA, PCIe topology, NUMA placement, and multi-GPU systems, validating GPU-to-GPU, GPU-to-network, and GPU-to-storage communication after repairs.
  • Coordinate replacement and RMA activities for GPUs, GPU trays, baseboards, NVSwitch components, system boards, and NICs.
  • Administer enterprise Linux distributions across the production fleet, troubleshooting boot failures, kernel issues, systemd services, file systems, and OS performance.
  • Analyze system and kernel logs (journalctl, dmesg, lspci, dmidecode, ipmitool, ethtool, ss, sar, perf) and manage kernel modules, device drivers, DKMS packages, and OS updates.
  • Support Linux networking, storage mounts, authentication, and security controls, and investigate OS crashes, kernel panics, and out-of-memory conditions.
Server Hardware, Firmware, and Out-of-Band Management
  • Support enterprise server platforms from vendors such as Dell, HPE, ASUS, Supermicro, and NVIDIA, troubleshooting processors, memory, PCIe devices, NICs, storage controllers, and other hardware components.
  • Manage BIOS, BMC, CPLD, NIC, GPU, switch, and drive firmware, and perform remote console troubleshooting, power cycling, and log bundle generation.
  • Partner with on-site Data Center technicians on component replacement, cabling validation, break-fix activity, and post-repair testing.
High-Performance Networking
  • Troubleshoot Ethernet, Infini Band, and RoCE connectivity within HPC and AI environments, including link state, VLAN/bonding issues, MTU mismatches, packet loss, and RDMA/GPUDirect RDMA…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary