Lead AI Engineer; Inference Serving & Performance
Lead AI Engineer (Inference Serving & Performance)
Full time ·
01. ABOUT THE COMPANYDizzaract is a product-driven company operating at the intersection of gaming, digital platforms, and AI. We build and scale multiple products, including FAR Labs & Gamed — each exploring a different space, yet united by a shared approach: moving fast, staying curious, and focusing on things that people actually use.
We operate as a collaborative, non-hierarchical team where ideas are valued based on their impact, not their origin, and where AI is embedded across everything we build, from infrastructure to product decisions.
02. ABOUTTHE ROLE
FAR Labs is building a distributed inference platform designed to serve large language models efficiently across diverse hardware. Serving quality — latency, throughput, memory efficiency, cost, and model quality — sits at the core of the product.
We’re looking for a Lead AI Engineer — Inference Serving & Performance to own the technical direction of our inference stack and how we measure its performance.
You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to make our serving infrastructure faster and more efficient while maintaining a clear quality bar.
This is the senior technical seat for AI serving at FAR Labs. You will set the serving and measurement roadmap, establish engineering standards, and guide engineers working across serving, metrics, and testing.
The role is highly hands‑on. You’ll be expected to work directly with the inference stack, identify performance bottlenecks, implement improvements, and demonstrate their impact through rigorous and reproducible measurement.
03. WHO YOU ARE- Inference Expert: You have deep hands-on experience with LLM inference serving and understand what happens between a model being loaded and a production-grade token being served.
- Performance Engineer: You think naturally in terms of latency, throughput, memory bandwidth, GPU utilisation, model quality, and cost — and understand how changes across the stack affect each of them.
- GPU‑Native: You are comfortable working close to the hardware with CUDA or Triton and understand modern GPU architectures, memory behaviour, and precis ions such as FP8 and FP4.
- Distributed Systems Thinker: You understand the challenges of routing, scheduling, communication, caching, and workload distribution across heterogeneous compute infrastructure.
- Measurement‑Driven: You care about proving performance improvements properly. You understand load generation, latency percentiles, throughput‑at‑SLO, reproducibility, and the importance of measuring speed alongside quality.
- Technical Leader: You have previously led an inference, serving, or similarly complex technical initiative and can set direction that strong engineers can execute against.
- Builder: You enjoy solving technically difficult problems in environments where the architecture, tooling, benchmarks, and engineering standards are still evolving.
- Inference Serving:
Own and improve the FAR Labs inference serving stack, including prefill/decode disaggregation, continuous batching, KV-cache management, cache‑aware routing, speculative decoding, and other serving optimisations. - Serving Performance:
Improve throughput, latency, memory efficiency, and serving cost while maintaining clearly defined model‑quality standards. - GPU Performance Engineering:
Identify and eliminate GPU performance bottlenecks across the inference stack using CUDA, Triton, profiling, kernel optimisation, memory optimisation, and appropriate precision strategies. - Mixture‑of‑Experts Serving:
Drive efficient serving of large MoE models, including expert parallelism, all‑to‑all communication, expert‑load…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).