Senior AI Data Scientist — Agentic Process Development
Listed on 2026-09-13
-
IT/Tech
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Company Overview
team.blue is the market leader in enabling digital success for small and medium-sized businesses (SMBs) across Europe, catering to over 3 million customers in 25+ languages. Our mission is to make online business success simpler, by providing our customers with all the tools and resources they need to excel online and remain ahead of the curve.
Position OverviewWe are looking for a Senior AI Data Scientist to streamline HR processes e — not by analysing them, but by building agentic systems to run them. Recruitment, onboarding, performance, rewards and offboarding are each multi-step processes spanning several systems and up to 25 countries, and your mandate would be to create systems that can streamline them end to end.
The method matters more than the domain: map a process, quantify what it costs in headcount, score which steps an agent could take, build a proof of concept, and take it to production. This work sits closer to building autonomous, side-effecting systems than to building predictive models. The agents you design would be able to revoke IT access, issue signed contracts, and flag pay outliers into approval workflows.
A wrong output here is not a bad number someone can catch — it is a high impact action taken in the world.
Not whether you can hand‑roll a gradient‑boosted tree. LLM coding tools can do that faster than you can. Classical ML and applied statistics are the entry fee for this role — necessary, and assumed. Everyone we are talking to has them.
What separates candidates is whether you can build an agent that is robust, cost‑effective and trustworthy — with deterministic operations rather than “LLM does everything” patterns. Building a demo is now easy. Knowing whether to trust one is not.
We also mean end to end literally. You write it, you containerise it, you instrument it, and you own it when it breaks.
Your day would involve- Time with the HR Ops lead mapping how a leaver actually gets offboarded across 19 countries — then turning that into a process inventory with FTE cost attached per step
- Facilitating a half‑day session with Talent, Rewards and HR Ops leads to score automation candidates on impact, feasibility and LLM/tool fit — extracting requirements live from people who do not think in data models
- Designing the state transitions: what triggers, what branches, which systems get called, where it waits, when it escalates, and what happens when step 4 of 9 fails
- Building the guardrails before the capability — dry‑run mode, an approval gate ahead of anything irreversible, least‑privilege scoped credentials, a rollback path
- Deciding where a human stays in the loop, at what confidence threshold, and designing a review queue they will actually use
- Writing evals for output that precision and recall do not capture: task‑completion rate, hallucination rate, gendered or culturally biased language in AI‑drafted reviews
- Wiring an agent to a webhook instead of a nightly batch pull — and making the handler idempotent so a retry does not offboard someone twice
- Deciding which steps in a flow warrant a frontier model and which can run on something cheap, then proving that routing decision with numbers
- Sitting in a vendor demo asking what their API actually exposes, what their data model looks like, and what integration really costs us
- 7+ years building data and ML systems in industry, spanning both sides of the LLM shift. We want the judgment that comes from having debugged systems before you could ask a model what was wrong.
- Somewhere in that history: you have shipped something that had permission to take an irreversible action affecting real customers — and you can tell us what you did to sleep at night.
- Expert in…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).