Sr. Data Center Technician; GA,
Listed on 2026-09-25
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software)
At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.
Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.
Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.
At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.
JobOverview
We are looking for an experienced Senior Onsite Data Center Professional to apply specialized technical expertise and advanced judgment in a fast-paced, complex AI/HPC environment. Rather than merely executing established procedures, this role demands independent analysis, complex problem-solving, and decisive action to address large-scale infrastructure challenges. You will take full ownership of technical solutions and strategic recommendations, serving as a primary authority for issue resolution rather than simply escalating problems.
In this senior capacity, you will drive process optimizations, provide critical technical guidance to cross-functional teams, and actively shape operational strategies to support cloud-scale compute and storage environments. Adaptability, advanced technical judgment, and the ability to influence outcomes are central to your success in this role.
This position will be onsite in Columbus, Georgia at the customer’s data center.
Responsibilities- Advanced Diagnostics & Problem Solving:
Exercise independent technical judgment to lead advanced hardware diagnostics, isolate root causes, and own the resolution of complex AI/HPC issues (e.g., GPU kernel hangs, interconnect anomalies) ensuring optimal cluster stability. - Technical Authority & Ownership:
Serve as the definitive technical authority and primary escalation point within the client ticketing system, taking full ownership of complex hardware and network challenges to develop solutions rather than merely escalating them. - Continuous Process Improvement:
Drive operational efficiencies by evaluating current workflows and implementing systemic process improvements. Lead the documentation and continuous refinement of operational strategies and Standard Operating Procedures (SOPs). - Technical Guidance & Mentorship:
Provide expert technical guidance, mentorship, and training to junior staff on complex ticket resolution, physical interventions, and safety protocols, heavily influencing team development and decisions. - Risk Assessment & RCA:
Identify risks, assess systemic impacts, and lead cross-functional Root Cause Analysis (RCA) investigations for recurrent hardware or facility failures, recommending and implementing…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).