More jobs:
Principal Infrastructure Engineer, AI Cluster & Validation
Job in
Houston, Harris County, Texas, 77246, USA
Listed on 2026-09-09
Listing for:
Socket.dev
Full Time
position Listed on 2026-09-09
Job specializations:
-
IT/Tech
AI Engineer (Applied/Software)
Job Description & How to Apply Below
Overview
As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters.
Your work will directly translate into higher system availability, compute optimization and reduced operational costs.
- Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
- Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerators to validate cluster behavior under genuine load rather than synthetic proxies alone, and translate what those runs reveal into fleet-wide fixes.
- Lead deep diagnosis of large-scale cluster failures and performance regressions
, isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, Infini Band/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior. Serve as the final escalation point for the hardest slow-job and stalled-job investigations. - Design and build the validation and burn-in systems that qualify nodes, racks, and full pods at scale — NCCL/RCCL collective sweeps, HPL/HPCG and MLPerf-style benchmarks, thermal and power soak tests, straggler and flapping-link detection — and automate them so that qualification is a repeatable pipeline, not a manual campaign.
- Drive cluster optimization end to end
, tuning fabric configuration (adaptive routing, QoS and congestion control, SHARP in-network reduction, rail and topology-aware placement), collective communication libraries and algorithm selection, GPUDirect RDMA and storage paths, and host-level settings (huge pages, IRQ affinity, CPU governors, MIG and driver configuration) to convert raw hardware into delivered throughput. - Partner with Infrastructure, Platform, SRE, and customer-facing teams to translate operational and customer performance needs into durable engineering solutions, and to feed diagnostic signal back into provisioning, remediation, and capacity workflows.
- Build production-grade Python systems and performance tooling for automated triage, telemetry correlation, and regression detection, leveraging AI tools to accelerate delivery. Assess impact to the team's software and validation stack from new hardware product programs, and explore AI-driven process improvement and automation.
- Establish engineering standards for reliability, observability, benchmarking methodology, and operational excellence across all services, and raise the diagnostic capability of the wider organization through mentorship, runbooks, and post-incident technical write-ups.
- Education: Bachelor's or higher degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
- Experience: 10+ years of relevant experience building, operating, or debugging large-scale compute infrastructure, including significant time at staff or principal level owning cross-team technical direction.
- AI Workload Expertise: Hands-on experience running real AI compute jobs at scale — pre-training, fine-tuning, or large-scale inference of open-source or proprietary models — including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×