×
Register Here to Apply for Jobs or Post Jobs. X

AI Research Engineer, Model Optimization and Inference at Iconic Interactive

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Jack & Jill
Full Time position
Listed on 2026-07-30
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Game & 3D/XR Development, Software Engineer
Salary/Wage Range or Industry Benchmark: 70000 - 110000 GBP Yearly GBP 70000.00 110000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Job Title

AI Research Engineer, Model Optimization and Inference

Salary

Not Disclosed

Company Description

Iconic Interactive is a seed-stage London-based AI-native video game studio building intelligent virtual actors that speak, move, and react in real-time. The team includes researchers and engineers from leading AI labs and game studios developing character intelligence, narrators, and world directors to create personal, immersive entertainment universes.

Job Description

As an AI Research Engineer at Iconic Interactive, you will bridge the gap between massive research models and real-time interactive entertainment. You'll architect high-performance inference pipelines for multimodal LLMs and TTS models, optimizing them to run on consumer hardware. Your work ensures digital entities perform seamlessly within the constraints of a game loop.

Location

London, UK

Why this role is remarkable
  • Lead the technical frontier by taking cutting-edge research models and making them run in real-time on consumer hardware for interactive digital experiences.
  • Join a seed-stage startup founded by AI lab and game studio veterans, offering significant autonomy and end-to-end ownership of the inference engine.
  • Work at the unique intersection of System ML and Game Tech, shaping the core intelligence of virtual actors in next-generation entertainment.
What You Will Do
  • Architect and maintain low-latency inference pipelines for Multimodal LLMs, TTS, and Vision models targeting server-side and consumer edge environments.
  • Implement state-of-the‑art optimization techniques like Speculative Decoding, KV-Cache quantization, and custom CUDA/Triton kernels to minimize latency and maximize throughput.
  • Collaborate with game engineering teams to integrate thread-safe, non‑blocking asynchronous inference directly into the game loop using C++ wrappers.
The ideal candidate
  • Holds an MSc or PhD in Computer Science or ML with deep expertise in model optimization techniques like quantization, pruning, and distillation.
  • Possesses strong proficiency in C/C++ and Python, with hands‑on experience deploying latency‑sensitive ML models using frameworks like PyTorch or JAX.
  • Demonstrates specialized knowledge in LLM‑specific optimizations such as KV-cache management, speculative decoding, and high‑performance runtimes like TensorRT or ONNX.
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary