×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer; LLM

Job in Belfast, County Antrim, BT1, Northern Ireland, UK
Listing for: Ocho
Full Time position
Listed on 2026-09-02
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 100000 GBP Yearly GBP 100000.00 YEAR
Job Description & How to Apply Below
Position: Machine Learning Engineer (LLM)

Machine Learning Engineer (LLM & Agent Evaluation)

  • Mid to senior ML Engineering role at a fast-scaling, AI-first software business
  • Build scalable evaluation infrastructure and multi-agent architectures at the core of the product
  • Belfast based, hybrid with async-friendly global team
  • Salary: competitive, reflecting experience, with equity
  • UK work authorisation required

About the Company

Our client is a fast-scaling AI software business powering enterprise automation for Fortune 500 clients including major names across financial services and healthcare technology. Their small, elite Data Science and AI team builds and deploys cutting-edge ML and agentic AI systems at scale, with a culture built around intellectual curiosity, hands-on leadership and pragmatic startup thinking. Leaders stay close to the code, debate ideas openly and move fast without corporate inertia.

This is a team where exceptional engineers thrive.

The Role

A newly created individual contributor position for an ML Engineer who wants to own evaluation end to end. You will design robust evaluation frameworks, build automated scoring and regression testing pipelines, and track quality across model, prompt and agent behaviour changes over time. A core part of this role involves building the infrastructure that converts expensive frontier agent tokens into optimised internal neural inference, a genuine and proprietary competitive advantage.

Working closely with engineering teams and data scientists, you will analyse edge-case failure modes, build real-time quality dashboards and ensure high-confidence deployment workflows as the system scales.

Key Responsibilities

  • Design evaluation frameworks and metrics covering accuracy, safety, latency and cost across agent and LLM systems
  • Build automated scoring pipelines, rubric-based grading and LLM-as-judge systems that scale beyond manual review
  • Design and build chained multi-agent architectures that underpin the core product capability
  • Stand up automated regression suites that catch quality drops from model, prompt or agent-logic changes before they reach production
  • Bridge high-cost frontier models into low-cost, optimised internal inference systems
  • Build dashboards and reporting that track model and agent quality over time and across releases
  • Identify failure modes and edge cases across diverse scenarios, working with the team to prioritise fixes
  • Partner closely with engineers building agent capabilities and with the senior data scientist for deeper analytical support

What You'll Need

Essential:

  • Bachelor's degree in Computer Science, Machine Learning, Statistics or a related field, or equivalent practical experience
  • 3 or more years of experience in ML engineering, NLP or applied data science with hands-on exposure to LLM or agent-based systems
  • Deep practical experience building, chaining and evaluating autonomous agent workflows and frontier LLMs
  • Practical experience building or operating evaluation frameworks, automated scoring or benchmark systems for ML and LLM outputs
  • Strong Python skills and comfort building data pipelines for evaluation datasets
  • Solid understanding of NLP and modern LLM capabilities including prompting techniques, agentic workflows and retrieval
  • Working experience with Google Cloud Platform including Vertex AI and Big Query, or equivalent AWS or Azure experience
  • Experience with inference optimisation and bridging frontier models into lower-cost internal systems

Desirable:

  • Experience with LLM-as-judge techniques, rubric design or human-in-the-loop evaluation programmes
  • Familiarity with agent architectures and the specific failure modes of multi-step and agentic systems
  • Experience operating evaluation systems at scale in a production environment

Why Apply?

  • Competitive salary plus equity, with package tailored to UK, NI or European candidates
  • Work on genuinely novel evaluation and agent infrastructure at the competitive core of a scaling AI product
  • Hands-on, intellectually driven team that debates ideas openly and challenges assumptions constructively
  • Async-friendly culture with approximately three syncs per week, designed to protect deep focus and work-life balance
  • Fast-moving startup environment with full operational autonomy and no corporate inertia
  • Belfast based with a globally distributed, elite AI engineering team behind you

Interested?

For a confidential conversation about this opportunity, connect with Justin Donaldson on Linked In or submit your CV via the link below.

Skills:
ML NLP Python Cloud ML Ops

Benefits:
Equity

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary