×
Register Here to Apply for Jobs or Post Jobs. X

Lead AI Engineer; Inference Serving & Performance

Job in Abu Dhabi, UAE/Dubai
Listing for: Dizzaract
Full Time position
Listed on 2026-09-15
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 420000 - 720000 AED Yearly AED 420000.00 720000.00 YEAR
Job Description & How to Apply Below
Position: Lead AI Engineer (Inference Serving & Performance)

Lead AI Engineer (Inference Serving & Performance)

Full time ·

01. ABOUT THE COMPANY

Dizzaract is a product-driven company operating at the intersection of gaming, digital platforms, and AI. We build and scale multiple products, including FAR Labs & Gamed — each exploring a different space, yet united by a shared approach: moving fast, staying curious, and focusing on things that people actually use.

We operate as a collaborative, non-hierarchical team where ideas are valued based on their impact, not their origin, and where AI is embedded across everything we build, from infrastructure to product decisions.

02. ABOUT

THE ROLE

FAR Labs is building a distributed inference platform designed to serve large language models efficiently across diverse hardware. Serving quality — latency, throughput, memory efficiency, cost, and model quality — sits at the core of the product.

We’re looking for a Lead AI Engineer — Inference Serving & Performance to own the technical direction of our inference stack and how we measure its performance.

You will work across inference serving, GPU performance, distributed systems, benchmarking, and model optimisation to make our serving infrastructure faster and more efficient while maintaining a clear quality bar.

This is the senior technical seat for AI serving at FAR Labs. You will set the serving and measurement roadmap, establish engineering standards, and guide engineers working across serving, metrics, and testing.

The role is highly hands‑on. You’ll be expected to work directly with the inference stack, identify performance bottlenecks, implement improvements, and demonstrate their impact through rigorous and reproducible measurement.

03. WHO YOU ARE
  • Inference Expert: You have deep hands-on experience with LLM inference serving and understand what happens between a model being loaded and a production-grade token being served.
  • Performance Engineer: You think naturally in terms of latency, throughput, memory bandwidth, GPU utilisation, model quality, and cost — and understand how changes across the stack affect each of them.
  • GPU‑Native: You are comfortable working close to the hardware with CUDA or Triton and understand modern GPU architectures, memory behaviour, and precis ions such as FP8 and FP4.
  • Distributed Systems Thinker: You understand the challenges of routing, scheduling, communication, caching, and workload distribution across heterogeneous compute infrastructure.
  • Measurement‑Driven: You care about proving performance improvements properly. You understand load generation, latency percentiles, throughput‑at‑SLO, reproducibility, and the importance of measuring speed alongside quality.
  • Technical Leader: You have previously led an inference, serving, or similarly complex technical initiative and can set direction that strong engineers can execute against.
  • Builder: You enjoy solving technically difficult problems in environments where the architecture, tooling, benchmarks, and engineering standards are still evolving.
04. RESPONSIBILITIES
  • Inference Serving:
    Own and improve the FAR Labs inference serving stack, including prefill/decode disaggregation, continuous batching, KV-cache management, cache‑aware routing, speculative decoding, and other serving optimisations.
  • Serving Performance:
    Improve throughput, latency, memory efficiency, and serving cost while maintaining clearly defined model‑quality standards.
  • GPU Performance Engineering:
    Identify and eliminate GPU performance bottlenecks across the inference stack using CUDA, Triton, profiling, kernel optimisation, memory optimisation, and appropriate precision strategies.
  • Mixture‑of‑Experts Serving:
    Drive efficient serving of large MoE models, including expert parallelism, all‑to‑all communication, expert‑load…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary