Senior ML Inference Architect GenAI Systems
Listed on 2026-07-27
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Architect, Cloud Engineer - Software
Description
AWS Neuron is the software stack powering AWS Inferentia and Trainium machine learning accelerators, designed to deliver high-performance, low-cost inference Neuron Serving team develops infrastructure to serve modern machine learning models—including large language models (LLMs) and multimodal workloads—reliably and efficiently on AWS silicon. We are seeking a Software Development Engineer to lead and architect our next-generation model serving infrastructure, with a particular focus on large-scale generative AI applications.
Description
AWS Neuron is the software stack powering AWS Inferentia and Trainium machine learning accelerators, designed to deliver high-performance, low-cost inference Neuron Serving team develops infrastructure to serve modern machine learning models—including large language models (LLMs) and multimodal workloads—reliably and efficiently on AWS silicon. We are seeking a Software Development Engineer to lead and architect our next-generation model serving infrastructure, with a particular focus on large-scale generative AI applications.
Key job responsibilities
- Architect and lead the design of distributed ML serving systems optimized for generative AI workloads
- Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem
- Design and implement scalable solutions for both offline and online inference workloads
- Lead integration efforts with frameworks such as vLLM, SGLang, Torch XLA, TensorRT, and Triton
- Develop and optimize system components for tensor/data parallelism and disaggregated serving
- Implement and optimize custom PyTorch operators and NKI kernels
- Mentor team members and provide technical leadership across multiple work streams
- Drive architectural decisions that impact the entire Neuron serving stack
- Collaborate with customers, product owners, and engineering teams to define technical strategy
- Author technical documentation, design proposals, and architectural guidelines
You'll Lead Critical Technical Initiatives While Mentoring Team Members. You'll Collaborate With Cross-functional Teams Of Applied Scientists, System Engineers, And Product Managers To Architect And Deliver State-of-the-art Inference Capabilities. Your Day Might Involve
- Leading design reviews and architectural discussions
- Rapidly prototyping software to show customer value
- Debugging complex performance issues across the stack
- Mentoring junior engineers on system design and optimization
- Collaborating with research teams on new ML serving capabilities
- Driving technical decisions that shape the future of Neuron's inference stack
The Neuron Serving team is at the forefront of scalable and resilient AI infrastructure focus on developing model-agnostic inference innovations, including disaggregated serving, distributed KV cache management, CPU offloading, and container-native solutions. Our team is dedicated to up streaming Neuron SDK contributions to the open-source community, enhancing performance and scalability for AI workloads. We're committed to pushing the boundaries of what's possible in large-scale ML serving.
Recent Shares
https://awsdocs
- Basic Qualifications
- 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- 5+ years of non-internship professional software development experience
- Experience as a mentor, tech lead or leading an engineering team
Preferred Qualifications
- Master's degree in computer science or equivalent
- Deep expertise in ML Frameworks/Libraries such as JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT.
Los Angeles County applicants:
Job duties for this position include: work safely and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).