×
Regístrese Aquí para solicitar empleo o publicarlo X

Senior Data Scientist — AI Evaluation & Quality | Internal AI Agents; Remote in Europe

Online/Remoto - Ideal para candidatos en
08001, Barcelona, Cataluna, España
Empresa: Lever, Inc.
Remoto/Desde casa puesto
Publicado en 2026-10-04
Especializaciones laborales:
  • TI/Tecnología
    Ingeniero de IA, Evaluación de IA, Inteligencia Artificial
Rango Salarial o Referencia de la Industria: 70000 - 110000 EUR Anual EUR 70000.00 110000.00 YEAR
Descripción del trabajo
Puesto: Senior Data Scientist — AI Evaluation & Quality | Internal AI Agents (Remote in Europe)
About Finom

Finom is a European tech startup headquartered in Amsterdam, and we’re on a journey towards revolutionizing the financial landscape for entrepreneurs worldwide. Our mission is to develop an all-in-one financial B2B solution that integrates banking functions, accounting, financial management, and invoicing into a seamless, mobile-first platform.

We recently closed a €115 million Series C equity round (around $133 million), bringing our total funding to approximately $346 million. This significant investment follows a $105 million growth funding round from General Catalyst, a long-term backer since 2021 known for supporting companies like Airbnb, Hub Spot, KAYAK, and Stripe.

Finom's platform goes beyond traditional banking, offering invoicing and a growing suite of features, including AI-enabled accounting, aiming to simplify financial management for entrepreneurs. We're actively expanding our reach across key EU markets like Germany, France, the Netherlands, Italy, and Spain.

At Finom, we’re not just redefining the entrepreneurial experience — we’re empowering our employees to make a real difference. Your work matters, and your impact extends far beyond product metrics. We nurture innovation and an inspiring work environment where bold ideas thrive, prioritizing thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and our business as a whole.

Maintaining our start-up spirit, we prioritize thorough research, swift implementation of solutions, and ensuring that every effort we make benefits our users, employees, partners, and, of course, our business.

AI Team

You’ll join Finom’s AI Team as the founding IC dedicated to the quality, evaluation and telemetry of AI agents powering Finom's internal operations and tools (including Ops workflows and internal AI analytics engines across ~20 core processes).

Our belief

An AI agent is only as good as the evaluation loop running on it. Because internal operational agents directly touch financial, compliance and support workflows, evaluation is an exact statistical and engineering discipline here.

Your mission

Design the evaluation methodology, build golden benchmarks, and establish quality gates for our internal AI agents from scratch—working directly with process owners and domain experts, with no senior quality owner above you to lean on.

Core Stack

Databricks, Deep Eval, Claude Code, Cursor, Python, SQL, dbt.

What You Will Be Doing
  • Own and extend our offline/online evaluation suites across ~20 internal AI agent processes—datasets (capability + regression), LLM-as-a-judge rubrics, and deterministic checks.
  • Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt, context, or tool changes hit production.
  • Work directly with domain experts to label cases and resolve annotator disagreement—fixing definition criteria rather than averaging disagreement away.
  • Build test datasets derived from real user & operational traffic (tickets, internal chats, colleague queries) rather than synthetic edge cases.
  • Harden statistical methodology: handle judge drift, verbosity bias, non-determinism, and measure true metric shifts vs. noise.
  • Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. cost/latency trade-offs.
Must-Haves
  • 5+ years in Data Science / Product Analytics / Applied AI roles, with sustained product-level metric ownership.
  • Production LLM

    Experience:

    In the last 1–2 years, you have built, shipped, or evaluated LLM-based systems (RAG, multi-step tool use, agents) as a core, primary job responsibility.
  • Autonomous Quality Ownership:
    Proven track record of owning evaluation…
Requisitos del puesto
10+ años Experiencia laboral
Para ver y solicitar empleos que acepten solicitudes de su ubicación o país, toque el botón a continuación para realizar una búsqueda.
(Si este trabajo está en su jurisdicción, entonces puede estar usando un Proxy o VPN para acceder a este sitio, para seguir avanzando, debe cambiar su conectividad a otro dispositivo móvil o PC).
 
 
 
Busque más trabajos aquí:
(Ingrese pocas palabras para obtener mejores resultados)
Localización
Aumentar el radio de búsqueda (millas)
0
200
Filtros
Nivel Educativo
Experiencia mínima requerida (años)
Publicado en los últimos:
Salario