Modeling Architect
Listed on 2026-09-10
-
Engineering
Systems Engineer, Hardware Engineer
About Neurophos
The demand for new data centers and AI compute is rapidly outpacing the planet's energy capacity. Digital solutions are hitting a power wall as we approach the physical limits of traditional silicon. Conquering this bottleneck means rethinking the fundamental architecture of inference compute. The industry's current path can't meet the need, so we're taking a different approach.
Instead of traditional electronic circuits, we use silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. Our optical cells are 10,000x smaller than traditional photonic components, enabling unprecedented density. By using photonics instead of electricity, our chips become more efficient as they scale. This architecture will deliver up to 100 times the energy efficiency of existing solutions while significantly improving performance for large-scale AI inference.
We’ve assembled a world-class team of industry veterans and recently raised a $110M Series A led by Gates Frontier. Participants include M12 (Microsoft’s Venture Fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, and others.
Join us and shape the future of computing!
Location: Austin, TX or Sunnyvale, CA. Full-time onsite position.
Reports To
:
Sr. Director of Modeling
FLSA Status
:
Exempt
We are seeking a modeling architect for hands‑on architecture modeling of the T100 optical inference accelerator, with hardware/software co‑design in the loop. You will work alongside senior engineers across two tracks. The first is analytical and system performance: roofline and limiter analyses, architecture performance models, workload setup, and the resulting plots and reports. The second is hardware models: functional, performance, and power models of compute blocks, memory, and hardware/software interfaces built in an event‑driven simulator, plus RTL simulation.
We expect real depth in one track and will help you build breadth across both. Most engineers start with a bounded piece, a single workload, a hardware block, or one layer of the model stack, and take on the surrounding area as the models mature.
Key ResponsibilitiesBring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, hybrid/SSM models, prefill versus decode, KV cache, expert routing, and quantization, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator.
Bind Hugging Face and PyTorch workloads to the programming model and run them on the functional model.
Co‑design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, and multi‑chip mapping.
Build in one or more layers of the modeling stack: roofline and limiter studies;
Python energy and latency models; C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM; cycle‑approximate performance and power models; and RTL simulation with Verilator and System Verilog.Own the tests, configs, and plots behind a result so anyone can rerun it and see what was assumed.
Use coding agents on real multi-file edits, and own the review of the C++ and System Verilog they generate.
Share results with the architects setting the design, as well as the compiler, runtime, and RTL teams, and carry their questions back into the model stack.
BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
3+ years of experience in hardware modeling, performance simulation, computer architecture, or related work.
Proficiency in Python or modern C++ (C++17 or later). Python-first and C++-first backgrounds are both welcome.
Working…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).