×
Register Here to Apply for Jobs or Post Jobs. X

Technical Director, -Scale AI Model Inferencing

Job in San Jose, Santa Clara County, California, 95199, USA
Listing for: Conductor
Full Time position
Listed on 2026-09-12
Job specializations:
  • Software Development
    Software Architect, Backend Developer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 219000 - 351000 USD Yearly USD 219000.00 351000.00 YEAR
Job Description & How to Apply Below

San Jose, California, United States

Please Note:

To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period.

Advancing the World’s Technology Together

Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future.

We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities.

What You’ll Do

Inference is becoming a memory‑bandwidth business. As models scale past what any single GPU can hold — KV caches grow with context, MoE expert weights spill beyond HBM, and new architectures change the rules of what “model state” even means — the winners will be the companies that treat memory as the core product of AI inference
, not an afterthought.

We are looking for a Hands‑on Principal Engineer who combines deep, first‑principles knowledge of AI model architectures (dense Transformers, Mixture‑of‑Experts, State Space Models, and hybrids) with production‑scale inference expertise
, to own the requirement for full‑stack AI memory solutions at scale — spanning GPU HBM, host DRAM, CXL‑attached memory pools, and NVMe/SSD tiers and Samsung Cognos, AI memory software that moves model state intelligently across them.

This person will be the technical authority who connects model behavior to memory‑system design: someone who can explain why an MoE router's activation pattern dictates an LRU expert cache policy, why a Mamba state cache breaks the assumptions of Paged Attention, and why disaggregated prefill/decode changes the required memory bandwidth per token by an order of magnitude — and then build the products that exploit those facts.

Location: Daily onsite presence at our San Jose office/headquarters in alignment with our Flexible Work policy

Job : 43027

Model Architecture Expertise — The Foundation
  • Serve as expert on how different model families consume and move memory, and translate that into memory-product requirements:
    • Dense Transformers
      : MHA/MQA/GQA/MLA attention, KV-cache growth characteristics, long-context behaviors, attention sinks and prefix locality.
    • Mixture‑of‑Experts
      : routed vs. shared experts, expert‑parallel execution, routing skew and hot‑expert locality, expert‑weight offloading and cache‑admission policies, per‑token weight‑read economics.
    • State Space Models (Mamba/Mamba‑2) and hybrid SSM‑attention architectures
      : recurrent state vs. KV cache semantics, state size per sequence and per layer, cache‑swapping behavior for context switching and batching, and what “cache‑aware scheduling” means when the state is a fixed‑size tensor instead of a token‑indexed table.
    • Emerging architectures
      : linear attention, sliding‑window/hybrid layers, diffusion and multimodal transformers — and how each changes the memory hierarchy math.
  • Model the memory footprint, bandwidth demand, and access patterns of frontier open‑weight models (e.g., Llama/Qwen‑class dense, Deep Seek/Kimi‑class MoE, Jamba‑class hybrids) and publish internal reference architectures for each.
  • Track the model landscape as a roadmap input: anticipate what coming architectures (longer contexts, agentic multi‑session reuse, reasoning‑loop workloads, speculative decoding drafts) will demand from memory systems 12–24 months out.
Large‑Scale Inference Expertise
  • Own deep expertise in production inference stacks…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary