Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Listed on 2026-07-18
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit that accelerates deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK is the backbone for accelerating deep learning and GenAI workloads, offering an ML compiler, runtime, and application framework that integrates with popular ML frameworks like PyTorch and JAX to deliver top‑performance inference and training.
The Inference Enablement and Acceleration team focuses on running a wide range of models and supporting new architectures while maximizing performance on AWS’s custom ML accelerators. Working across the stack from PyTorch to the hardware‑software boundary, engineers build systematic infrastructure, innovate new methods, and create high‑performance kernels, ensuring every compute unit is fine‑tuned for optimal performance. This role offers the chance to work at the intersection of machine learning, high‑performance computing, and distributed architectures, shaping the future of AI acceleration technology.
Key responsibilities include leading distributed inference support for PyTorch in the Neuron SDK, tuning these models for highest performance on AWS Trainium and Inferentia silicon, developing low‑level optimizations, and collaborating with compiler, runtime, framework, and hardware teams to optimize machine learning workloads for the global customer base.
Key job responsibilities- Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
- Participate in all stages of the ML system development lifecycle – distributed computing architecture design, implementation, performance profiling, hardware‑specific optimizations, testing, and production deployment.
- Build infrastructure to systematically analyze and onboard multiple models with diverse architectures.
- Design and implement high‑performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.
- Analyze and optimize system‑level performance across multiple generations of Neuron hardware.
- Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.
- Implement optimizations such as fusion, sharding, tiling, and scheduling.
- Conduct comprehensive testing, including unit and end‑to‑end model testing with continuous deployment and releases through pipelines.
- Work directly with customers to enable and optimize their ML models on AWS accelerators.
- Collaborate across teams to develop innovative optimization techniques.
- Bachelor's degree in computer science or equivalent.
- 5+ years of professional software development experience.
- 5+ years of design or architecture experience for new and existing systems.
- Fundamentals of machine learning and large‑language models, including architecture, training, and inference life cycles, and experience optimizing model execution.
- Software development experience in C++ and Python (experience in at least one language required).
- Strong understanding of system performance, memory management, and parallel computing principles.
- Proficiency in debugging, profiling, and implementing best software engineering practices in large‑scale systems.
- Familiarity with PyTorch, JIT compilation, and AOT tracing.
- Familiarity with CUDA kernels or equivalent low‑level kernels.
- Experience with performant kernel development such as CUTLASS or Flash Infer.
- Knowledge of syntax and tile‑level semantics similar to Triton.
- Experience with online/offline inference serving using vLLM, SGLang, TensorRT, or similar platforms in production.
- Deep understanding of computer architecture, operating systems, and parallel computing.
The Inference Enablement and Acceleration team fosters a builder’s culture where experimentation is encouraged, impact is measurable, and collaboration, technical ownership, and continuous learning are valued. The team provides mentorship, thorough but kind code reviews, and projects that develop engineering expertise, empowering members to handle complex tasks.
ResourcesLearn more about Neuron:
-
https://(Use the "Apply for this Job" box below).-success
Equal OpportunityAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).