AI Research Engineer, Model Optimization and Inference at Iconic Interactive
Listed on 2026-07-30
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Game & 3D/XR Development, Software Engineer
Job Title
AI Research Engineer, Model Optimization and Inference
SalaryNot Disclosed
Company DescriptionIconic Interactive is a seed-stage London-based AI-native video game studio building intelligent virtual actors that speak, move, and react in real-time. The team includes researchers and engineers from leading AI labs and game studios developing character intelligence, narrators, and world directors to create personal, immersive entertainment universes.
Job DescriptionAs an AI Research Engineer at Iconic Interactive, you will bridge the gap between massive research models and real-time interactive entertainment. You'll architect high-performance inference pipelines for multimodal LLMs and TTS models, optimizing them to run on consumer hardware. Your work ensures digital entities perform seamlessly within the constraints of a game loop.
LocationLondon, UK
Why this role is remarkable- Lead the technical frontier by taking cutting-edge research models and making them run in real-time on consumer hardware for interactive digital experiences.
- Join a seed-stage startup founded by AI lab and game studio veterans, offering significant autonomy and end-to-end ownership of the inference engine.
- Work at the unique intersection of System ML and Game Tech, shaping the core intelligence of virtual actors in next-generation entertainment.
- Architect and maintain low-latency inference pipelines for Multimodal LLMs, TTS, and Vision models targeting server-side and consumer edge environments.
- Implement state-of-the‑art optimization techniques like Speculative Decoding, KV-Cache quantization, and custom CUDA/Triton kernels to minimize latency and maximize throughput.
- Collaborate with game engineering teams to integrate thread-safe, non‑blocking asynchronous inference directly into the game loop using C++ wrappers.
- Holds an MSc or PhD in Computer Science or ML with deep expertise in model optimization techniques like quantization, pruning, and distillation.
- Possesses strong proficiency in C/C++ and Python, with hands‑on experience deploying latency‑sensitive ML models using frameworks like PyTorch or JAX.
- Demonstrates specialized knowledge in LLM‑specific optimizations such as KV-cache management, speculative decoding, and high‑performance runtimes like TensorRT or ONNX.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: