Data Center Technician L2
Job in
Reno, Washoe County, Nevada, 89550, USA
Listed on 2026-09-11
Listing for:
Covestic Inc
Full Time
position Listed on 2026-09-11
Job specializations:
-
IT/Tech
IT Infrastructure, Systems Engineer, Systems Administrator, Unix/Linux
Job Description & How to Apply Below
The Data Center Operations Technician II - GPU Specialist is responsible for supporting and maintaining highly available GPU-based compute environments, engineering labs, and data center infrastructure. This role partners closely with hardware, software, QA, and systems engineering teams to deploy, troubleshoot, and optimize next-generation computing platforms. The ideal candidate combines strong data center operations experience with a deep understanding of GPU technologies, server hardware, Linux/Windows administration, and large-scale test infrastructure.
Key Responsibilities Compute Farm & Infrastructure Operations- Manage and maintain a high-performance compute farm consisting of builders, packagers, testers, and supporting infrastructure.
- Monitor system health, availability, and performance to ensure operational excellence and SLA compliance.
- Lead system recovery efforts and incident response activities to minimize downtime and restore services quickly.
- Support deployment, configuration, and lifecycle management of GPU servers, workstations, and test systems.
- Perform rack, stack, cabling, hardware installation, and equipment decommissioning activities within the data center.
- Collaborate closely with system architects, hardware engineers, software engineers, QA teams, and platform operations teams to develop, test, debug, and release next-generation products.
- Troubleshoot hardware, software, networking, and infrastructure issues impacting engineering and validation environments.
- Provide technical support for GPU systems, PCBs, servers, storage systems, and network-connected devices.
- Assist engineering teams with validation, benchmarking, and deployment activities for new technologies and platforms.
- Gather operational metrics and performance data to identify trends, risks, and improvement opportunities.
- Develop, maintain, and enhance Standard Operating Procedures (SOPs), runbooks, and technical documentation.
- Drive continuous improvement initiatives that increase availability, throughput, operational efficiency, and test accuracy.
- Participate in change management activities and ensure documentation is kept current.
- Support and troubleshoot Linux, Windows, and macOS environments.
- Utilize scripting and automation tools to streamline operational tasks and improve scalability.
- Maintain accurate asset and infrastructure records using DCIM systems.
- Assist with infrastructure automation and configuration management initiatives.
- Associate's degree or Bachelor's degree in Engineering, Information Technology, Computer Science, or a related technical field; equivalent experience will be considered.
- 5+ years of experience supporting data center operations, engineering labs, high-performance computing environments, or related technical infrastructure.
- Experience working with GPU-based systems, PCBs, servers, and large-scale system deployments.
- Proficiency with DCIM platforms such as Nautobot or similar infrastructure management tools.
- Experience with scripting and automation technologies including Shell, Python, and Ansible.
- Working knowledge of networking fundamentals and protocols including: TCP/IP DNS NFS SSL/TLS
- Experience administering and troubleshooting:
Linux Windows macOS - Strong troubleshooting and problem-solving skills across hardware, operating systems, networking, and infrastructure.
- Excellent written and verbal communication skills with the ability to present technical concepts to non-technical audiences.
- Strong teamwork skills and the ability to work effectively in cross-functional engineering environments.
- Experience managing High Performance Computing (HPC) environments.
- Experience utilizing cluster management and workload scheduling platforms such as:
Bright Cluster Manager (BCM) Slurm - Industry certifications such as CCNA or equivalent networking certifications.
- Advanced Windows and Linux systems administration experience.
- Understanding of modern data center architecture, including:
Compute infrastructure Storage platforms Networking systems - Knowledge of data center facilities infrastructure with emphasis on liquid-cooled environments.
- Experience supporting AI, machine learning, or GPU-intensive workloads.
- Strong mechanical aptitude and comfort performing hands-on hardware installation, maintenance, and repair tasks.
- Data Center Operations
- GPU Infrastructure Management
- Linux & Windows…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×