×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Senior Linux Administration Santa Clara, CA; Onsite

Job in Santa Clara, Santa Clara County, California, 95050, USA
Listing for: E-Solutions
Full Time position
Listed on 2026-08-30
Job specializations:
  • IT/Tech
    Unix/Linux
Job Description & How to Apply Below
Position: Senior Linux Administration :- Santa Clara, CA(Onsite)

Senior Linux Administration

Location:

Santa Clara, CA (Onsite)

AI and HPC Infrastructure

Position Description

The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.

What This Candidate Will Be Doing

  • Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
  • Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, file system, service, and resource-consumption issues across bare-metal server fleets.
  • Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
  • Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
  • Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
  • Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
  • Automate repeatable administration and remediation tasks with Bash and Python.
  • Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.

What We Need To See

  • 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
  • Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
  • Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
  • Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
  • Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
  • Strong shell scripting and Python-based automation capability.
  • Working knowledge of storage and network dependencies affecting Linux host health.
  • Ability to operate independently in ambiguous, high-severity production situations.

Preferred Experience

  • Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
  • Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.
  • Experience supporting validation labs or pre-production cluster certification.
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary