Software Engineer, AI Inference Runtime
Listed on 2026-09-12
-
Software Development
AI Engineer (Applied/Software), Software Engineer, Machine Learning/ ML Engineer, DevOps
As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models.
You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm’s AI platform.
Responsibilities:- Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms.
- Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency.
- Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration.
- Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.
- Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices.
- 5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference.
- Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior.
- Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement.
- Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware.
- Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths.
- Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution.
- Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization.
- Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries.
You will be part of our AI Platforms team – A driven and diverse group passionate about developing foundational production capabilities to support AI inference provide a collaborative setting where your ideas can come to life quickly. The success of our AI projects will be directly influenced by your work, crafting the company’s AI inference capabilities and defining and operating production inference workloads.
This is an outstanding opportunity to work with world-class teams and contribute to groundbreaking advances in AI technology. Join us in building the next generation of AI inference infrastructure!
$209,100-$282,900 per year
We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm's offering. The total reward package will be shared with candidates during the recruitment and selection process.
Accommodations at ArmAt Arm, we want to build extraordinary teams. If you need an adjustment or an…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).