Principal Modeling Architect
Listed on 2026-09-01
-
Software Development
AI Engineer (Applied/Software)
OXMIQ designs GPU and AI silicon for large-scale model inference and training, and is building the system software and analysis platforms that prove out those designs before they reach customers. We are a startup driving innovation across the full stack —
from atoms to agents — and we move fast on the strength of our own tools.
The Principal Performance Modeling Architect owns OXMIQ's system-solution performance analysis — the modeling platform (OxSol) and its Speed-of-Light and OxFabric analysis capabilities that convert a proposed silicon and cluster architecture into simulation-backed numbers: performance, energy, and cost-per-token across heterogeneous accelerator topologies. These numbers feed our architecture decisions, our pricing, and our customer proposals directly, and they are produced before we commit capital — which makes this one of the highest-leverage technical roles at the company.
This is a hands-on individual-contributor role
. You own the platform's technical direction in partnership with engineering leadership, and you write the hard parts yourself. The platform models LLM inference today; your mandate is to extend its reach to new workload classes and make the engine more extensible as it grows.
We are looking for someone who has seen it and done it — a senior engineer with a tinkering mindset who is energized by being the person who can answer "is this architecture actually optimal, and what does a token cost on it?" and would rather prove it with a defensible model than hand-wave it. You read simulation output critically, reason about whether an outcome is physically sensible, hold a high bar for code quality, and treat AI-assisted development as a standard part of how you work.
Key Responsibilities- Own and extend OXMIQ's system-solution performance modeling platform end to end — Speed-of-Light modeling of performance, energy, and cost-per-token across heterogeneous accelerator hardware and cluster topologies.
- Extend the platform beyond LLM inference to training workloads and vision-transformer / multimodal models
, defining the modeling approach for each new workload class. - Influence silicon and system architecture decisions by turning proposed designs into defensible, simulation-backed numbers ahead of capital commitment.
- Make the platform more extensible and agent-driven, including exposing analysis capabilities through the Ox Capsule interface.
- Validate and reason about simulation results — correlating model outputs against the observed behavior of real workload frameworks (e.g., vLLM, SGLang, training stacks) and continuously tightening model fidelity.
- Set and own the platform's technical direction in line with the executive vision, uphold engineering and code-quality standards, and provide technical guidance to interns supporting the work.
- Serve as a technical point of contact in selected proposal and partner engagements.
- 5–8 years of relevant experience in performance modeling, computer architecture, or systems performance engineering, with a demonstrable, shipped body of work
. - Hands-on experience building analytical / first-principles / roofline performance models of AI or HPC systems — deriving models, not only running benchmarks.
- Solid understanding of both inference and training workloads and how they stress compute, memory bandwidth, and interconnect.
- Strong working knowledge of AI accelerator and cluster hardware
: GPUs/XPUs, HBM and memory hierarchies, scale-up/scale-out fabrics, and how parallelism strategies (tensor/pipeline/expert/data) map onto them. - Proficiency in Python for building and maintaining a production-quality codebase — clean abstractions, tests, and the discipline to keep a large platform maintainable.
- A track record of influencing architecture decisions and a tinkering mindset — comfortable getting into datasheets, specs, and code to find out why a number is what it is.
- An AI-first mindset
: fluent use of AI-assisted development workflows (Claude Code or equivalent), and the judgment to review and reason critically about simulation results, outcomes, and code quality
. - Strong written and verbal communication; able to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).