×
Register Here to Apply for Jobs or Post Jobs. X

Full-Stack AI Compute Architect

Job in Campbell, Santa Clara County, California, 95011, USA
Listing for: Oxmiq Labs
Full Time position
Listed on 2026-07-09
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Software Architect, Software Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 150000 - 200000 USD Yearly USD 150000.00 200000.00 YEAR
Job Description & How to Apply Below

Updated Role | Now hiring:
Full-Stack AI Compute Architect About OXMIQ

OXMIQ provides complete hardware and software GPU IP that lets our customers build their own AI silicon. Founded by Raja Koduri, we recently closed a $35M Series A (co-led by Samsung Catalyst Fund and Fundomo, with Media Tek, Intel Capital, and others), bringing total funding to $60M, with Jim Keller joining our board.

Our OxCore architecture combines scalar, tensor, SIMT, and orchestration engines on one platform — delivering CUDA compatibility alongside a fully programmable architecture. On top of it runs OxPython, our software stack that runs existing CUDA and PyTorch code unmodified with day-zero support for new AI models. We have an established team and a working stack, and we’re now shifting from development to production — hardening the platform and scaling it for inference at real customer scale.

You’ll join the OXMIQ architecture team, which sets the direction of our software, silicon, and systems. You’ll help refine our software strategy and deliver the highest-performance solutions by optimizing the whole stack — from AI models and frameworks at the top down to kernel and hardware-level code generation.

The Scope of This Role

This role spans the AI software stack from models on down — the full vertical path a workload travels: from a model authored in frameworks such as PyTorch, JAX, or ONNX, through graph capture and the compiler, into the runtime and scheduler, down to individual kernels and the instructions that execute on OxCore.

We’re looking for an architect who is fluent across that whole range — someone who can reason about how a transformer or diffusion model is structured and about warp divergence and register allocation, and who can make the choices that connect those layers into one coherent, high-performance platform. You won’t be the deepest specialist at every layer, but you should be able to hold the whole range in your head and drive decisions end to end.

Key Responsibilities
  • Refine the technical direction of OxPython, the OXMIQ software stack — from framework front-ends (PyTorch, JAX, ONNX) and model ingestion, through graph capture, the compiler, runtime, and driver layers, down to kernels — in close partnership with the existing software team.
  • Own how AI models map onto the platform: understand how modern workloads (LLMs, diffusion, vision, agentic inference, and beyond) are structured and expressed at the framework level, and shape how they are captured, partitioned, and lowered so they run unmodified with day-zero support and scale efficiently for inference.
  • Own whole-stack performance: drive optimization end to end, identify and remove bottlenecks at every layer — from graph-level and framework-level inefficiencies down through compiler, runtime, driver, and kernels — and close the gap to peak hardware utilization.
  • Guide the graph and compiler strategy: MLIR dialect and IR design, lowering pipelines, operator representation, and LLVM backend code generation targeting Oxmiq hardware IP.
  • Architect work orchestration across OxCore’s scalar, tensor, and SIMT engines, and shape the compiler/runtime approach that runs CUDA-optimized models unmodified — preserving day-zero support for new AI models without sacrificing programmability.
  • Advance the SIMT execution and parallel compute strategy: thread/warp scheduling, divergence handling, occupancy, and the memory hierarchy, and how parallel workloads are expressed, lowered, and mapped onto the hardware.
  • Refine and optimize the strategy for kernel programming models (including Triton-based and custom op lowering), kernel fusion, and tiling, in collaboration with kernel engineers.
  • Design for scale and production: ensure solutions grow cleanly with customer workloads, deployment sizes, and model complexity, and harden them for real-world delivery.
  • Work with architecture and hardware teams to translate performance requirements into efficient code generation and runtime strategies, and feed software needs back into hardware definition.
  • Drive the software/hardware validation strategy — evolving how the stack and silicon are co-verified as we scale from bring-up…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary