Systems Modeling Engineer
Listed on 2026-10-01
-
IT/Tech
Systems Engineer, AI Engineer (Applied/Software) -
Engineering
Systems Engineer, AI Engineer (Applied/Software)
Artificial intelligence (AI) is transforming our world. It can perform cognitive functions that previously only humans could do, such as perceiving interactions across different modalities and environments - with the ability to quickly learn and then solve complex problems. Tensordyne is an AI system solution company that builds very high-performance, low-power generative AI inference systems. Our mission, through the creation of custom silicon, hardware and software, is to enable multimodal Generative AI inference acceleration at scale, with safe, sustainable, high-performance systems for our hyperscaler and neocloud data center customers.
We are at the leading edge of advancing the latest research and product improvements for generative Al inference solutions that will make Al even more advantageous for compelling new generative AI applications. Tensordyne is a well funded, fast-paced startup company with headquarters in both Sunnyvale, CA, and Munich, Germany. We also have many talented team members working remotely across North America and Europe.
We take care of our people and their families with comprehensive benefits, competitive compensation, flexible spending options, and recognition programs, because building category-defining technology starts with a healthy, supported team. Come join us as we shape the future of multimodal generative artificial intelligence!
We are looking for a Systems Performance Modeling Engineer to build the models and tools that predict how generative AI inference workloads perform on Tensordyne systems, from a single accelerator up through rack, pod, and cluster scale. This is a hands-on engineering role for someone who likes writing simulator code, running experiments, and digging into why a prediction and a measurement don't match.
Working closely with our architects and the silicon, hardware, networking, and software teams, you'll capture workload behavior, extend simulation and analytical models of our silicon, interconnect, and multi-hop fabrics, and validate them against real hardware. You'll be comfortable moving across the stack, from the model graph through collectives to the network fabric, to track down where performance is going.
What You'll Do- Implement and extend simulation-based performance models for multimodal generative AI inference at rack, pod, and cluster scale, covering compute, memory, collective communication, and network fabric.
- Model how serving strategies (tensor, pipeline, and expert parallelism, prefill/decode disaggregation, batching, and KV-cache placement) interact with Tensordyne silicon and fabric topology, and measure the effect on latency, throughput, and cost per token.
- Build trace-capture and replay tooling that records real execution from our inference runtime and replays it under hypothetical silicon, system, and network configurations.
- Model collective communication on multi-hop scale-out fabrics, including implementing custom collective algorithms designed for our topology.
- Run calibration experiments on Tensordyne hardware as systems come up, compare them against model predictions, and fix the sources of error.
- Run design-space sweeps and write up clear analyses that architects and engineering teams use in ASIC, fabric, and system configuration decisions.
- Produce performance projections that support product and customer discussions.
- Keep the modeling codebase fast, tested, and reproducible so other engineers can run it themselves.
- Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture.
- Solid understanding of distributed ML execution, including parallelism strategies, collective communication…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).