×
Register Here to Apply for Jobs or Post Jobs. X

LLM Reliability Engineer

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: usepassionfruit.com
Full Time, Part Time position
Listed on 2026-07-27
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer, AI QA / Validation Engineer
Salary/Wage Range or Industry Benchmark: 70000 - 90000 GBP Yearly GBP 70000.00 90000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Passionfruit launches PIPAI Learn More Passionfruit careers

LLM Reliability Engineer

Hybrid Full-Time

About Passionfruit

Passionfruit is reshaping how marketing teams work. We started as a marketplace connecting brands with specialist marketers, and we're now building PIP, an AI-native platform that's changing how marketing operations actually get done.

We're not building generic AI tools and hoping marketers find them useful. We're building with marketing teams, because that's the only way to make something that genuinely works.

About PIP

PIP is our AI-native workspace for modern marketing operations. It centralises context, insights, and workflows into one intelligent platform - helping teams analyse performance, automate reporting, and extract meaningful answers from their data in minutes rather than days.

We ship fast, iterate based on real feedback, and scale deliberately. As we onboard larger enterprise customers and deepen LLM-driven workflows, reliability is the single biggest lever for trust and retention.

The Role

We're hiring an LLM Reliability Engineer to own the reliability layer of PIP - ensuring our AI systems consistently deliver useful, trustworthy, and production-ready experiences for users.

This role is focused on a newer and increasingly important challenge: understanding whether the AI is actually helping users achieve what they need, identifying where trust breaks down, and improving reliability before issues become churn.

You'll sit close to real user sessions, LLM traces, and production workflows - spotting silent failures, inconsistent outputs, frustrating responses, and workflows that stall. You'll work across observability, evaluations, and product feedback loops to improve the quality and resilience of our AI systems over time.

You'll be the person who notices that our top user this week was also our angriest user, and does something about it.

What You'll Do
  • Monitor LLM analytics, error tracking, and session replays (Post Hog, Arize, App Signal, Langfuse) to spot user frustration and silent failures before they're reported
  • Set up and tune sentiment analysis, error-clustering, anomaly alerts, and reliability dashboards so the team gets early signals on quality issues rather than learning about them weeks later
  • Run structured evaluations on AI outputs - consistency, accuracy, usefulness, hallucination rates, and task completion - across our agent library and core workflows
  • Build and maintain eval datasets and regression testing systems for prompts, retrieval pipelines, and agent behaviours
  • Partner with product and engineering to translate observed user friction into concrete prompt improvements, orchestration fixes, retrieval improvements, or product changes
  • Investigate edge cases and production failures across LLM workflows, identifying root causes and reliability gaps
  • Own a weekly view of platform reliability: what broke, what frustrated users, what trends are worsening, and what's actually been fixed
What We’re Looking For
  • 4+ years in a similar quality-focused role
  • A genuine instinct for edge cases - you find the things others didn't think to test
  • Comfort working with LLM-powered products and a strong appetite to go deeper on evals, tracing, observability, and reliability tooling
  • Familiarity with concepts like hallucination detection, prompt regression testing, response evaluation, and agent reliability
  • Strong attention to detail paired with the conciseness to communicate issues clearly to engineers
  • Ability to work independently - you'll help define what reliability means here - while collaborating closely with engineering, product, and CS
  • A product-focused mindset: you care about whether users trust and successfully use the system, not just whether outputs technically succeed
  • Comfortable in a fast-moving environment where requirements evolve weekly
Nice to Have
  • Hands-on experience with Post Hog, Arize, Langfuse, Lang Smith, Helicone, or similar LLM observability tools
  • Familiarity with prompt engineering, RAG systems, or agent orchestration frameworks
  • Prior software development experience (a plus, not mandatory)
  • Exposure to Elixir/Live View or similar
  • Background working with data platforms, analytics tools, or AI-native products
Working Hours and Location
  • Location:

    London, Chancery Lane (3 days per week in-office)
  • Hours:

    9–6pm
#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary