×
Register Here to Apply for Jobs or Post Jobs. X

Senior​/Software Engineer – LLM Inference & Reinforcement Learning Platform

Job in Bellevue, King County, Washington, 98009, USA
Listing for: Snowflake
Full Time position
Listed on 2026-09-15
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Salary/Wage Range or Industry Benchmark: 236000 - 309750 USD Yearly USD 236000.00 309750.00 YEAR
Job Description & How to Apply Below

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact.

We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization
.

Our mission is to build the next generation of high-performance and intelligent inference systems
. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.

Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.

Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering
, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development
.

Recent innovations from Snowflake AI Research include Arctic Inference
, our open-source inference system, and technologies such as Shift Parallelism
, which dynamically adapts parallelism to workload characteristics;
SwiftKV
, which reduces redundant prefill computation;
Arctic Speculator and Suffix Decoding for fast speculative decoding;
Jacobi Forcing for causal parallel decoding; and Semi-Persistence for fast model swapping and dynamic multi-model serving.

This is an exciting opportunity to collaborate with a world-class team, including founding members of Deep Speed, vLLM, and Tensor Flow. Together, we will push the boundaries of AI systems and bring cutting‑edge research into production‑scale AI.

Responsibilities
  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
  • Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.
  • Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search,…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary