More jobs:
HPC System Administrator Consultant
Job in
Baton Rouge, East Baton Rouge Parish, Louisiana, 70804, USA
Listed on 2026-05-29
Listing for:
Louisiana State University
Full Time
position Listed on 2026-05-29
Job specializations:
-
IT/Tech
IT Support, Systems Engineer, Systems Administrator, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
If you close the browser or exit your application prior to submitting, the application progress will be saved as a draft. You will be able to access and complete the application through "My Draft Applications" located on your Candidate Home page.
Job Posting
Title:
HPC System Administrator Consultant
Position Type:
Professional / Unclassified
Department:
LSUAM FA - ITS - TA - RETS - HPC - Systems (Timothy William Wright ))
Work Location:
0340 Fred
C. Frey Computing Services Building
Pay Grade:
Professional
Job Description:
This position is for a "hands-on" IT Consultant in the High Performance Computing group in the Information Technology Services Department HPC IT Consultant specializes in hardware architecture and advanced troubleshooting to support, optimize, and maintain research computing infrastructure. This role is responsible for enabling both existing and emerging high performance computing initiatives through direct technical support, training, and system hardware and software support.
This role is designed for a technical expert who is equally comfortable troubleshooting physical hardware in the data center as they are writing complex automation scripts in a Red Hat Enterprise Linux (RHEL) environment. All Information Technology Services employees are expected to demonstrate a commitment to exemplary customer service in all facets of their work.
Job Responsibilities:
Operations:
Expertise and leadership in verifying the quality of operations of Linux supercomputers, infrastructure systems, and other research computing systems. This includes, but is not limited to, performing daily system checks, analyzing system logs, troubleshooting hardware and software problems, monitoring and analyzing storage/infrastructure/job performance, helping users recognize job performance problems, writing scripts to enhance monitoring, and responding to unplanned system events such as power outages.
This may require travel to various HPC sites to maintain physical installation of systems located off-site.
Proactively perform hardware maintenance on the clusters, cluster infrastructure, and other systems as needed. This includes diagnosing and fixing problems which includes, but is not limited to, running diagnostics, re-seating dimms, replacing hard disks, calling vendors for RMA support, replacing mother boards, and return shipping replacement parts.
Plan and perform software maintenance on both the clusters and the cluster infrastructure as needed. This includes, but is not limited to, installing operating systems, installing security patches, installing or upgrading drivers, upgrading firmware, installing or upgrading software licenses, installing or upgrading software specific to HPC cluster management. (50%)
Research:
Investigates, architects and implements new technology as appropriate to add new features to both the user environment and to our deployment environment. This requires the ability to work without training to take a new technology through installation to production. This also includes the ability to develop and document procedures related to that technology and to train other members of the group.
(25%)
Customer Support:
Respond to tickets which include complaints, requests, troubleshooting, assessing storage options, etc. Provide training to groups or individuals as needed. (15%)
Other duties as assigned. (10%)
Minimum Qualifications:
Bachelor's Degree with 3 years of experience (Ph.D. in Computation Science, Engineering or other computationally intensive disciplines substitutes for 2 year exp).
Experience in IT systems administration in Linux/HPC environments. Strong knowledge of Linux/Unix operating systems. Expertise in scripting and programming in bash and other languages.
Experience with HPC cluster resource managers and other management software such as Kickstart, DNF, RACADM, SSH keys, Ansible, etc. Experience working with, managing, and repairing hardware in large complex HPC systems. Proven experience troubleshooting complex hardware, networking, and performance issues in Linux-based HPC environments.
+LSU values skills, experience, and expertise. Candidates who have relevant experience in key job responsibilities are encouraged to apply- a degree is not required as long as the candidate meets the required years of experience specified in the job description.
Preferred Qualifications:
Master's degree in Computation Science, Engineering or related computationally intensive disciplines. 5 years of experience in Linux system administration in a large HPC deployment. Experience in scripting and programming.
Experience with PBS torque and moab, Infini Band, and Lustre file systems. Experience working with computational research projects utilizing large and complex HPC systems. Scripting skills in bash or similar language.
Experience with file systems and hardware such as Lustre, GPFS, NAS, DDN, Panasas.
Experience…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×