Senior Data Scientist
Listed on 2026-09-03
-
IT/Tech
Data Scientist, AI Evaluation, AI Engineer (Applied/Software)
About Warden AI
Warden AI helps HR Tech vendors, staffing firms, and employers prove and improve the fairness, accuracy, and compliance of their AI hiring systems. In a fast-changing regulatory environment, organizations need a trusted partner to validate their AI, and Warden is quickly becoming the default solution.
In 2025, Warden grew from 4 to 40 customers. In 2026, we’ve already more than doubled ARR and are growing faster every quarter. That growth is driven not just by market demand, but by something rarer: a genuine network effect. Customers want to promote working with Warden because it helps them win their own deals.
About the roleWe're hiring a Senior Data Scientist / Senior Applied Scientist to run highly defensible bias audits of high-stakes AI systems, used everywhere from startups to industry-leading enterprises. These audits carry real weight with regulators, courts, customers and candidates, and the methodology behind them has to hold up. The role spans AI system evaluation, rigorous statistical analysis, synthetic data generation, and an applied understanding of hiring and selection procedures.
You’ll join a team with deep expertise in behavioral and social sciences, evaluation methodology, and structured bias testing. You’ll need strong statistical and probabilistic reasoning: fluency with techniques like uncertainty quantification, experimental design, causal inference and Bayesian modeling. Just as essential is commercial instinct: the judgment to design rigorous, workable methods within real product goals and constraints, built through meaningful, hands-on experience shipping at startup speed.
You will report to the CTO and work closely with the founders and product team across hands-on analysis, methodological design, and strategic thinking. As one of a small number of data hires, you will have high agency to shape both how our analytical function evolves and the scope of your own role as we grow.
What you’ll doHere are a few examples of things you might be working on:
- Set and uphold rigorous statistical methodology. Define the statistical tests, fairness metrics, sampling strategies, and evaluation frameworks we rely on, and embed the checks and validation patterns that keep our analytical work accurate, reproducible, and defensible.
- Translate regulations and standards into practical tests. Work with our policy experts to translate legal requirements, guidance, and emerging HR and AI standards into clear, practical audit methodologies.
- Design the foundations for audit execution. Create the datasets, test frameworks, workflows, and analysis patterns that enable consistent, efficient, and high-quality audits, and help establish the data partnerships that inform our synthetic data generation and audit methodology.
- Analyze and interpret audit results. Assess what results do and don't show given uncertainty and limitations, and shape how findings are presented to customers through our reports and dashboards, including actionable insight like pointing to the likely source when an audit fails.
- Support key high-stakes conversations. Bring technical authority on data, methodology, and statistical defensibility to stakeholder discussions, including with customers' own data science, legal and compliance teams.
- Document and defend our methodology. Write the accessible explanations, white papers and industry reports that let regulators, customers and outside experts scrutinize our approach and trust the results.
- 5+ years in a commercial, high-performing product organization doing applied statistics or data science - not solely academic or research-lab experience.
- Strong statistical and probabilistic reasoning
, with hands-on fluency in confidence intervals and hypothesis testing, power analysis, experimental design, causal inference, and Bayesian modeling. - Experience designing evaluation methodologies - including black-box and counterfactual approaches - and the test datasets behind them, with good judgment on realism and validity limits.
- Practical familiarity with evaluating AI systems and their failure modes, across classic ML models, LLM-based systems, and AI agents.
- Fluen…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: