×
Register Here to Apply for Jobs or Post Jobs. X

Modeling Architect

Job in Sunnyvale, Dallas County, Texas, 75182, USA
Listing for: Neurophos, Inc.
Full Time position
Listed on 2026-09-18
Job specializations:
  • Engineering
    Software Engineer, Systems Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

About Neurophos

The demand for new data centers and AI compute is rapidly outpacing the planet's energy capacity. Digital solutions are hitting a power wall as we approach the physical limits of traditional silicon. Conquering this bottleneck means rethinking the fundamental architecture of inference compute. The industry's current path can't meet the need, so we're taking a different approach.

Instead of traditional electronic circuits, we use silicon photonics and an active, programmable metasurface to perform matrix multiplications at the speed of light. Our optical cells are 10,000x smaller than traditional photonic components, enabling unprecedented density. By using photonics instead of electricity, our chips become more efficient as they scale. This architecture will deliver up to 100 times the energy efficiency of existing solutions while significantly improving performance for large-scale AI inference.

We've assembled a world-class team of industry veterans and recently raised a $110M Series A led by Gates Frontier. Participants include M12 (Microsoft's Venture Fund), Carbon Direct Capital, Aramco Ventures, Bosch Ventures, Tectonic Ventures, Space Capital, and others.

Join us and shape the future of computing!

Location: Austin, TX or Sunnyvale, CA. Full-time onsite position.

Reports To:
Sr. Director of Modeling

FLSA Status:
Exempt

Position Overview

We are seeking a staff-level modeling architect to build the path from a production model or application to two things: a performance and energy number Neurophos will stand behind, and a functional model that software can boot against before tape-out.

The T100 architecture is still moving, and the workloads are the models the industry is publishing now, so this is hardware/software co-design in practice. You will bind a workload to the programming model and runtime, run it on the model stack, and feed the result back into decisions on tiling, instruction set architecture (ISA), memory hierarchy, and multi-chip mapping. The team works between the principal architects, the RTL and physical design groups, and the compiler and runtime teams.

At this level, you own a workload or block area along with the methodology behind it, and you mentor the engineers building models in that area.

Key Responsibilities
  • Bring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, attention and KV cache, expert routing, quantization, and hybrid/SSM models, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator.

  • Bind Hugging Face and PyTorch workloads to the programming model and runtime, then run them on the functional model so that software and architecture are looking at the same behavior.

  • Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, network-on-chip (NoC) traffic, and multi-chip mapping across pipeline, tensor, and sequence parallelism, including collectives.

  • Run roofline and limiter analysis and design space exploration across microarchitecture options, resolving bottlenecks between the compiler view and the hardware.

  • Develop Python energy and latency models in Num Py, Pandas, and Matplotlib that cover operators, tiling, SRAM and HBM traffic, and optical GEMM and vector-unit time.

  • Implement bit-accurate C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM, including narrow arithmetic, so software can begin bring-up before tape-out.

  • Contribute to the C++ event-driven simulation kernel itself, including coroutines, timed components, and traces, rather than only calling into it.

  • Implement cycle-approximate and cycle-accurate performance, power, and area (PPA) models, and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary