×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer, Machine Learning Infrastructure

Job in London, Laurel County, Kentucky, 40741, USA
Listing for: Deliveroo
Full Time position
Listed on 2026-07-25
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Backend Developer
Salary/Wage Range or Industry Benchmark: 120000 - 190000 USD Yearly USD 120000.00 190000.00 YEAR
Job Description & How to Apply Below

Software Engineer, Machine Learning Infrastructure - Generative AI
About the Team

Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps Door Dash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and Deep Seek) ourselves - real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs - delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper).

We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

About the Role

You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and Door Dash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability.

This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.

You’re excited about this opportunity because you will…
  • Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
  • Work on our open-weights serving stack - real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) - alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
  • Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases
  • Push the cost and latency frontier of GPU inference - turning batch jobs that took days into hours and cutting inference cost by multiples - while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.
  • Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
  • Partner closely with ML engineers, product engineers, data scientists, and platform teams across Door Dash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
  • Shape the future of the centralized GenAI platform - including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques - enabling the next generation of AI-powered products, agents, automation, and personalization.
We’re excited about you because you have…
  • BSc, MSc, or PhD in Computer Science or equivalent
  • 3+ years of industry experience in software engineering
  • Strong backend engineering fundamentals, especially in Python and distributed systems.
  • Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
  • Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
  • Hands‑on experience with LLM inference and/or fine‑tuning of open‑weight models in production - serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine‑tuning (SFT/DPO/LoRA).
  • Ability to work across ambiguous, fast‑moving technical areas and turn customer use cases into reusable platform capabilities
  • Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary