Senior Infrastructure Support Engineer
Listed on 2026-07-13
-
IT/Tech
Systems Engineer, SRE/Site Reliability, IT Infrastructure, Network Engineer
About Nscale
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack—energy, data centres, GPU superclusters, orchestration, and AI services—delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
AboutThe Role (Job Purpose)
Senior Infrastructure Support Engineers are the senior technical escalation point within Infrastructure Support, owning the health of Nscale's GPU fleets and the high-performance fabrics that connect them. This is a hands‑on L2/L3 role operating at the intersection of GPU hardware, east‑west networking, Linux, and data centre operations—acting as the operational bridge between Support, DC Operations, and Engineering.
You Will- Own complex, ambiguous problems end-to-end and make decisive calls in a results‑driven environment, taking calculated risks where speed matters.
- Communicate technical detail clearly, specifically, and concisely— to engineers, to customers, and to leadership. We treat communication quality as a core engineering skill, not a soft skill.
- Influence without authority and build strong relationships with senior stakeholders across the business to get things done.
- Grasp new technical concepts quickly, stay curious, and know which questions to ask to get up to speed fast.
- Bring discipline and organisation: evidence‑led investigations, accurate records, clean handovers.
- Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes.
- Diagnose and remediate GPU node faults across the full stack—driver, firmware, and hardware layers—from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA.
- Own east‑west fabric health: run link-level diagnostics (mlxlink, ibdiagnet or equivalent), isolate transceiver, optics, cabling, and switch‑port faults, and validate topology across Infini Band and RoCE/high-speed Ethernet fabrics.
- Investigate data-path issues on high‑performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing.
- Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion.
- Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans.
- Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation.
- Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover.
- Design and implement automation scripts and small tools to reduce toil and human intervention.
- Act as a key escalation point for the Support Organisation, taking ownership of strategic decisions where results matter.
- Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews.
- Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion.
- Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise.
- Experience. 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands‑on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.
- Communication. Able to explain complex technical detail clearly, specifically, and concisely— in tickets, in incident updates, and face to face with customers and stakeholders at all levels.…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).