Member of Technical Staff - ClusterMAX
Listed on 2026-09-15
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software)
Employment Type:
Full-Time
Work Setting:
In-office/Remote
Work Location:
United States, New York, Mexico, San Francisco, Canada
Work Hours:
Office hours
Find out more here:
About Semi Analysis
Semi Analysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to state-of-the-art AI Models, CUDA kernels, and GPU cloud infrastructure. We are recognized as the leading authority on AI infrastructure, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.
We’re a global team of over 20 analysts & engineers, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually.
Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies.
Some of industry-shaping articles are:
Inference MAX:The world first open inference benchmark that continuous benchmarks performance of popular frontier models
MI300X vs H100 vs H200 Training
:
Extensive Benchmarking & Deep Dive into CUDA & ROCm training developer experience & performanceTrainium2 Architecture & Networking
:
Deep Dive into Amazon’s new Trainium2 rack scale system & 3D torus scale up topology
Find out more here:
The RoleCluster MAX™ is the industry-standard GPU cloud rating system — 95% market coverage by volume, 84 providers rated, 209 tracked, and 140+ customer interviews behind each release. We recently published our cluster TCO and goodput framework in How Much Do GPU Clusters Really Cost?, showing that goodput expense alone swings 6–21% of total cluster TCO depending on fault-tolerance approach. We are now testing providers for Cluster MAX 3.0 with expanded benchmarks, security requirements, and analysis.
As an MTS on Cluster MAX, you will build the benchmarks and run the evaluations that determine how the world's GPU clouds get rated.
Leading development of the next generation of Cluster MAX™ benchmarks: storage IO and bandwidth, NCCL/RCCL collectives, fault-tolerance and goodput measurement, multi-tenant security and isolation testing
Deploying to and evaluating dozens of GPU clusters across hyperscalers and neoclouds (GB300/GB200 NVL
72, B300/B200, H200, MI355X, TPUv7)Extending our TCO and goodput methodology (see the Cluster MAX TCO & Goodput calculator) into automated, reproducible tests
Working with executives and engineers at 200+ neoclouds, hyperscalers, and chip vendors
Authoring technical research analyzing benchmark results, reliability, security, and ease of use, with direct authorship recognition
Hands-on experience operating GPU clusters:
Slurm and/or Kubernetes, Infini Band/RoCE fabrics, distributed storageStrong Python and shell scripting; comfort building benchmark harnesses and CI pipelines
Understanding of distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery
Experience similar to our Technical Consultant profile is a plus: due diligence, TCO analysis, and client-facing technical communication
Security mindset: multi-tenant isolation, bare-metal vs virtualized trade-offs
Why Semi Analysis?
At Semi Analysis, you’re placed directly inside the most important conversations shaping the future of AI and semiconductors. You’ll develop a first-principles understanding of the technical and economic dynamics behind the world’s most consequential technology, with rare end-to-end visibility across the entire stack—from silicon and systems to models, software, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).