Research Engineer - AI-Optimized Inference
Listed on 2026-07-25
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer, Machine Learning/ ML Engineer, Software Engineer
Team:
Systems / AI Infrastructure
Location:
San Francisco (on-site)
AI Systems, Model Optimization
Infinity Artificial Intelligence Institute San Francisco Bay Area
Research Engineer - AI-Optimized InferenceInfinity Artificial Intelligence Institute San Francisco Bay Area
1 week ago 145 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Report this job
Company
:
Infinity
Team:
Systems / AI Infrastructure
Location:
San Francisco (on-site)
Type:
Full-time
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like Alpha Evolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing.
This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand‑tuned inference library on the market, not just match one once.
- The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
- The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec‑sheet number.
- The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
- The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
- Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
- Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.
- Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
- Familiarity with search- or evolution-based optimization. Alpha Evolve-style loops, super optimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
- A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
- Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
- Python and a systems language, plus real hands‑on kernel work…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).