×
Register Here to Apply for Jobs or Post Jobs. X

Principal Systems Engineer

Job in Houston, Harris County, Texas, 77001, USA
Listing for: Nscale
Full Time position
Listed on 2026-07-01
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 175000 - 225000 USD Yearly USD 175000.00 225000.00 YEAR
Job Description & How to Apply Below
Principal Systems Engineer GPU Supercluster Bringup  We are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute. As a startup, we move fast, operate with ownership, and expect technical leaders to define standardsnot just follow them.

The Role  We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites. You will own the full lifecycle of deploymentfrom rack design and fabric architecture to cluster validation frameworks and production readiness standards.

You will set the bar for performance, reliability, and operational excellence. This role combines deep hands-on expertise with system-level thinking and cross-functional leadership.
What You'll Do  End-to-End Supercluster Bringup Ownership   Define the technical standards for node, rack, and full-cluster bringup.
Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments).
Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads.
Establish cluster-level acceptance criteria and validation frameworks.
Performance & Fabric Architecture   Tune and validate NCCL, RDMA, GPUDirect, and collective operations ntify and eliminate performance bottlenecks across hardware, topology, and firmware layers.
Drive congestion control and fabric optimization strategies.
Define performance benchmarking methodology for AI training workloads.
Deployment Strategy & Scalability   Design repeatable deployment models for multi-site expansion.
Build automation frameworks for provisioning and cluster validation.
Establish deployment SLAs, quality gates, and operational readiness standards.
Reduce time-to-capacity while increasing reliability.
Technical Leadership   Serve as the escalation point for complex bringup and performance issues.
Mentor senior engineers and shape infrastructure best practices.
Influence hardware selection, rack topology, and data center design decisions.
Partner with executive leadership on infrastructure scaling strategy.
What We're Looking For  Required   10+ years of experience in large-scale infrastructure or HPC environments.
Proven experience bringing up large GPU clusters (hundreds+ GPUs).
Deep expertise in high-speed networking (Infini Band, RoCE, Ethernet fabrics).
Strong understanding of server architecture (PCIe, NUMA, memory hierarchy).
Experience debugging performance issues across compute and network layers.
Strong automation and systems-level thinking.
Strongly Preferred   Experience scaling AI training clusters for frontier models.

Experience with liquid cooling or ultra-high-density deployments.
Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF).
Experience defining infrastructure standards in a fast-growing organization.
What Success Looks Like   Superclusters are brought online quickly, predictably, and at peak performance.
Deployment processes scale from first cluster to multi-site expansion.
Infrastructure becomes a competitive advantage.
You define the technical blueprint for how we scale AI infrastructure.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$175,000 - $225,000 USD
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary