×
Register Here to Apply for Jobs or Post Jobs. X

Research Engineer - AI-Optimized Inference

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Infinity Artificial Intelligence Institute
Full Time position
Listed on 2026-07-23
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Software Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 260000 USD Yearly USD 180000.00 260000.00 YEAR
Job Description & How to Apply Below

Company :
Infinity ·
Team :
Systems / AI Infrastructure Location :
San Francisco (on‑site) ·
Type :
Full‑time ·
The Mission

The Mission

The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like Alpha Evolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.

This role points that loop directly  code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.

We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing.

This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand‑tuned inference library on the market, not just match one once.

What you’ll work on

You’ll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:

  • The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
  • The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip’s measured peak rather than its spec‑sheet number.
  • The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel’s individual peak and the stack’s actual delivered throughput is hiding.
  • The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
  • Correctness underneath all of it. A faster kernel that returns a different answer isn’t faster, it’s wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
  • Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.
What we’re looking for

We care about depth and range more than a checklist, but strong candidates will have most of the following:

  • Performance engineering on real inference or accelerator code. You’ve made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
  • Familiarity with search‑ or evolution‑based optimization. Alpha Evolve‑style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven’t yet.
  • A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
  • Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
  • Python and a systems language, plus real hands‑on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.
Nice to have
  • Written high‑performance attention or matmul kernels by hand, not just…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary