Member of Technical Staff - Frontier System Modelling
Listed on 2026-09-12
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software)
Employment Type :
Full-Time
Work Setting:
In-office/Remote
Work Location:
United States, New York, Mexico, San Francisco, Canada
Work Hours:
Office hours
Find out more here:
Semi Analysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to cutting-edge AI Models, software, and infrastructure. We are recognized as the leading authority on the semiconductor supply chain, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.
We’re a global team of over 50 analysts, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry shaping articles while participating in 40+ conferences annually.
Our newsletter reaches more than 200,000 subscribers worldwide, including senior management and c-suite leaders at the leading semiconductor and AI companies.
We also offer three core products:
Industry Models – we develop and publish industry models on accelerator shipments, data centre demand and supply, GPU total cost of ownership, and more. We work with hyperscalers, neoclouds, many of the world’s largest hedge funds, and government agencies.
Core Research – our public equity markets product, geared towards financial investors, distils our deep technical research and knowledge into key insights on technology and product trends.
Consulting and Technical Due Diligence – We conduct custom research and project work to guide key strategic and investment decisions for the largest private equity funds, leading venture capital firms, companies across the AI ecosystem, and government agencies.
We are looking for a highly motivated member of technical staff to join our engineering team to work on system modelling on 100k+ chip AI clusters. This is an unique opportunity to work on an high-visibility open source projects to architecture & model performance for the next generation of AI chips. If you’re passionate about performance engineering, system modelling, first principles, and want to work at the intersection of hardware and software, this is a rare chance to make industry wide impact.
As part of the interview process, you'll complete a paid coding challenge designed to reflect typical daily tasks petitive Compensation depending on experience, skillset, location & business needs
Develop an complex model in Python to predict performance (MFU, tok/s/gpu, tok/s/user) on current & future AI chip projects across both frontier LLM training & inference models
Implement modern parallelism & inference strategies to generate Pareto frontier roofline curves for performance on each chip
Develop microbenchmarks across multiple vendors (AMD, NVIDIA, TPU, Trainium) to calibrate the system model to as close to reality as possible
Build BoM estimates to model performance per total cost of ownership and performance per watt
Strong skills in Python
Deep understanding about disaggregated prefill, wide expert parallelism, tensor parallelism, pipeline parallelism, sequence parallelism, etc.
Experience with frontier MoE training and inference workloads
Knowledgeable about GEMM sizes & proportions of wall time & FLOPs on every major operator in an modern transformer
Intuition about Amdahl principles, arithmetic intensity, weak & strong scaling
Previous experience as an ML engineer or kernel programmer
Develop deep expertise in large-scale system modelling for AI infrastructure, including performance prediction across hyperscale (100k+ chip) clusters
Gain a strong first-principles understanding of how hardware,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).