Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
Job in
Cupertino, Santa Clara County, California, 95014, USA
Listed on 2026-07-18
Listing for:
Amazon
Full Time
position Listed on 2026-07-18
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Job Description & How to Apply Below
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
We are building AWS Neuron, an SDK that accelerates deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. As part of the Acceleration Kernel Library team, you will craft high‑performance kernels for ML functions, optimizing performance across hardware, compiler, runtime, and framework layers.
Key Responsibilities- Design and implement high‑performance compute kernels for ML operations on Neuron hardware
- Analyze and optimize kernel‑level performance across multiple generations of Neuron accelerators
- Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
- Implement compiler optimizations such as fusion, sharding, tiling, and scheduling
- Work directly with customers to enable and optimize their ML models on AWS accelerators
- Collaborate across teams to develop innovative kernel optimization techniques
- 5+ years of non‑internship professional software development experience
- 5+ years of programming in at least one software programming language
- 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, scaling)
- 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations
- Experience as a mentor, tech lead, or leading an engineering team
- Bachelor’s degree in computer science or equivalent
- 6+ years of full software development experience
- Expertise in accelerator architectures for ML or HPC (GPUs, CPUs, FPGAs, or custom)
- Experience with GPU kernel optimization and GPGPU computing (CUDA, NKI, Triton, OpenCL, SYCL, or ROCm)
- Demonstrated experience with NVIDIA PTX and/or AMD GPU ISA
- Experience developing high‑performance libraries for HPC applications
- Proficiency in low‑level performance optimization for GPUs
- Experience with LLVM/MLIR backend development for GPUs
- Knowledge of ML frameworks (PyTorch, Tensor Flow) and their GPU backends
- Experience with parallel programming and optimization techniques
- Understanding of GPU memory hierarchies and optimization strategies
Amazon is an equal‑opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Base salary range
: $ – $ annually (location‑specific).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×