Research Scientist; AI Behaviours
Listed on 2026-07-22
-
Research/Development
AI Business & Operations, Data Scientist
Research Scientist Opportunity
White Circle is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn't do. We automatically test, enforce, and continuously improve these policies at scale.
We've raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, Hugging Face, Mistral, Deep Mind, Datadog, Sentry, and others. We process over 100M+ API calls every month. We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model.
We're a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you're the one we need.
Your Responsibilities- Own research projects end to end – from an unclear concern to a falsifiable experiment, clean baselines, and a result you can defend.
- Develop automated audit agents that discover and characterize suspect model behavior at scale.
- Study how misalignment and bias actually show up when real users interact with agents, and turn what you find into evals our products can ship.
- Pressure-test frontier agents in realistic, high-stakes scenarios to find where they break before our customers do.
- Run white-box and black-box investigations to understand how AI models fail.
- Publish what you learn as public blog posts and conference papers, and feed the rest back into our internal guardrails.
- Have a track record of empirical research in agent behavior, model evaluation, alignment, or a closely adjacent area.
- Have strong ML engineering. You can independently build a research MVP involving fine-tuning, agent inference, and evals, without waiting on a platform team.
- Have evidenced skills in experimental design under real conditions: isolating agent failure modes, calibrating judges and baselines, and distinguishing genuine signal from artifact.
- Can take a vague behavioral question and define the experiment that answers it, when there's no playbook – then run it fast and iterate.
- Are an AI power-user – fluent with frontier models and coding agents in your daily work.
- Published research at A
* venues (NeurIPS / ICML / ICLR / ACL and similar). - Interpretability depth – familiarity with modern interp tooling and concepts (NLAs, SAEs, persona vectors, etc.) and the ability to run whitebox investigations on our internal and open-source models.
- An MSc or PhD in machine learning, computer science, cognitive science, computational neuroscience, physics, or a related quantitative field.
- AI safety fellowship (MATS, ASTRA, Anthropic Fellows, etc.), or a comparable self-directed research record.
- Paid time off in line with your local regulations, no matter where you work from.
- Work from Paris (hybrid) with a relocation package available, or work from London (note: we are currently unable to provide relocation support or medical insurance for London-based roles).
- Comprehensive medical insurance for our France-based team.
- All the hardware, tools, and services you need.
- Covered subscriptions for AI agents and IDEs.
- Team off-sites twice a year: we've recently been to the Alps and to Saint-Tropez.
Please submit your application in English.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).