Engineering
Listed on 2026-07-22
-
Software Development
AI Engineer (Applied/Software), Software Architect, Software Engineer
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, Open Evidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital.
Join us and help build the platform engineers turn to to ship AI products.
We're looking for an Engineering Manager to lead our GPU Kernel Engineering team, the group responsible for writing the low-level CUDA code that makes Baseten's inference stack faster than anyone else's. This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers.
You'll own the technical direction of a team working at the intersection of GPU architecture, ML systems, and production inference. Your engineers write CUDA kernels for GEMMs, attention mechanisms, and MoE routing, optimize at the warp and tensor-core level, and ship improvements that directly reduce latency and cost for the AI companies running their most critical workloads on Baseten.
This role is not for someone who wants to step away from the technical work. You'll be close enough to the code to credibly review it, set direction, and unblock your team, while also building the processes, culture, and roadmap that let a world-class kernel team operate at its best.
EXAMPLE INITIATIVESYour team owns work like:
- Baseten Embeddings Inference:
The fastest embeddings solution available - The Baseten Inference Stack
- Driving model performance optimization
- Lead, grow, and mentor a team of GPU kernel engineers; own hiring, performance, and career development
- Set technical direction for the kernel roadmap, balancing short-term inference wins with long-term architectural investments
- Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy
- Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams
- Establish and maintain a high technical bar for kernel quality, performance, and correctness across the team's output
- Review kernel designs and implementations with enough depth to give meaningful feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies
- Guide the team's approach to profiling and bottleneck identification using tools like Nsight Systems, Nsight Compute, and Torch Profiler
- Stay current on the NVIDIA GPU ecosystem (Hopper, Blackwell, and beyond) and translate architectural advancements into team priorities
- Build the processes that allow a highly technical, distributed team to ship with velocity and rigor
- Represent the kernel team's work to senior leadership and external audiences including industry conferences
- Contribute to Baseten's open-source GPU library presence and technical brand
- Proven experience leading a team of GPU or ML systems engineers, with a track record of hiring and developing strong technical talent
- Deep personal background in GPU kernel engineering. You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level
- Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology
- Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem
- Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership
- Hands-on experience with Triton, CUTLASS, or CuTe DSL
- Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing
- Open-source contributions to GPU libraries or inference frameworks
- Experience presenting technical work at NVIDIA…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).