×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer - GPU Kernel Authoring & Optimization

Job in Bellevue, King County, Washington, 98009, USA
Listing for: CoreWeave
Full Time position
Listed on 2026-07-23
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 182000 - 242000 USD Yearly USD 182000.00 242000.00 YEAR
Job Description & How to Apply Below

About The Role

Core Weave is the top-rated AI‑cloud for high-performance GPU infrastructure across AI/ML, visual effects, rendering, and real‑time inference. Our stack is engineered for speed, scale, and cost‑efficiency—an unmatched alternative to traditional hyperscalers.

We’re looking for a Senior Engineer for Core Weave’s Benchmarking & Performance team, focused on kernel authoring and optimization. You will write, profile, and tune the GPU kernels that sit on the critical path of large‑scale model serving—squeezing maximum throughput and minimum latency out of every SM, tensor core, and byte of memory bandwidth. You will also aid us in achieving industry‑leading end‑to‑end performance benchmarking publications such as MLPerf.

Responsibilities
  • Author, profile, and optimize CUDA kernels—GEMMs, attention, MoE routing, quantization, KV‑cache, and fused epilogues—on the critical path of LLM inference.
  • Optimize for the hardware: exploit tensor cores and tune occupancy, memory coalescing, shared‑memory/register usage, and overlap of compute with data movement.
  • Use kernel‑authoring DSLs and compilers to prototype and ship kernels quickly without sacrificing performance.
  • Benchmark rigorously: build reproducible microbenchmarks and roofline analyses, and validate that kernel‑level wins translate to end‑to‑end latency/throughput gains across model‑serving stacks (vLLM, TensorRT‑LLM, llm‑d, SGLang).
  • Implement and maintain benchmarking workflows for end‑to‑end MLPerf Inference (and Training) runs, including workload setup, cluster configuration, runbooks, and result validation.
  • Lead design reviews and drive architecture within the team; decompose multi‑service work into clear milestones.
  • Mentor junior engineers; review cross‑team designs and elevate coding/testing standards.
  • Help ensure reproducible, well‑documented benchmarking and kernel‑optimization processes.
Qualifications
  • 5+ years of experience building high‑performance computing, GPU/accelerator software, or performance‑critical systems.
  • Hands‑on CUDA experience is required—you have written and optimized custom kernels and are fluent with the CUDA programming and memory model.
  • Deep understanding of GPU architecture and performance: tensor cores, warp/occupancy tuning, the memory hierarchy and bandwidth, NVLink/PCIe, and profiling with Nsight Compute/Systems.
  • Strong coding in C++ and Python; comfortable reading and writing low‑level, performance‑sensitive code.
  • Familiarity with model‑serving stacks (vLLM, TensorRT‑LLM, llm‑d, SGLang) and the kernels that dominate their inference cost.
  • Strong communicator comfortable collaborating with cross‑functional teams and external partners.
Preferred
  • Triton or Mojo for authoring custom GPU kernels — highly desired.
  • CuTe DSL for Python‑based kernel authoring on NVIDIA GPUs.
  • JAX and its Pallas kernel language for authoring kernels on GPU/TPU.
  • HIP / ROCm and AMD GPU experience.
  • NCCL and collective‑communication performance.
  • Experience with alternative accelerators such as Google TPUs and Meta’s MTIA.
  • Familiarity with kernel‑authoring DSLs and nano‑compilers such as KNYFE and its Block DSL.
  • Experience with Kubernetes at production scale.
  • Experience with SUNK (Slurm on Kubernetes) / Slurm for scheduling large GPU jobs.
  • Experience running MLPerf submissions or similar large‑scale audited benchmarks.
  • Contributions to OSS projects such as vLLM, SGLang, PyTorch, Triton, or CUTLASS.
Wondering if you’re a good fit?

We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren’t a 100 % skill or experience match.

Why Core Weave?

Help shape an industry‑defining inference platform that enables teams to deploy generative AI and real‑time applications  squeezing every last microsecond out of GPU kernels and delivering reliable model serving excites you, this is the place to build. We’re in an exciting stage of hyper‑growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning.

Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

  • Be Curious at Your Core
  • Act…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary