More jobs:
Senior Kernel Engineer
Job in
Bellevue, King County, Washington, 98009, USA
Listed on 2026-09-20
Listing for:
Akraya, Inc.
Full Time
position Listed on 2026-09-20
Job specializations:
-
Software Development
Software Engineer, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Primary
Skills:
- CUDA programming (advanced)
- Triton optimization (advanced)
- Profiling tools (advanced)
- Debugging kernels (advanced)
- Performance tuning (advanced)
Contract Type: W2
Duration: 3+ months with possible extension or conversion
Location: Bellevue, WA (#LI-Onsite)
Pay Range: $85.00 - $90.00 Per hour on W2 #LP
Job SummaryThis role involves the low-level development and optimization of compute kernels for AI accelerator hardware. You will ensure training environments accurately mirror real-world kernel engineering challenges by writing, debugging, and optimizing operators. A key aspect of this position is using profiling tools and accelerator programming interfaces to validate task realism and define solution quality.
Key Responsibilities- Develop and optimize custom compute kernels for AI accelerators.
- Validate training environment tasks for realistic kernel development scenarios.
- Debug and profile kernel code to identify and resolve performance issues.
- Analyze kernel performance against hardware constraints and memory hierarchy.
- Consult with training environment development teams on task realism and quality.
- Computer Science, Computer Engineering, Electrical Engineering, or equivalent technical degree and 7+ years of relevant experience is required.
- Direct experience with AWS Trainium/Inferentia and the Neuron Kernel Interface (NKI) and Hands-on proficiency with a low-level kernel programming language such as CUDA, Triton, or the AWS Neuron Kernel Interface (NKI).
- Familiarity with the Neuron SDK, Neuron Profiler, and framework integration (PyTorch, JAX).
- Experience porting kernels between accelerator backends (e.g., CUDA/Triton to NKI) and understanding of deep learning operator implementations (attention, GEMM, normalization).
Experience in developing and optimizing low-level code for AI accelerators is essential, with a strong preference for experience with Triton over CUDA.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×