Evaluation Research Manager
Listed on 2026-09-12
-
Research/Development
Research Scientist, Research Analyst
About Aaru
Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes. Building a useful simulation requires more than generating plausible text.
Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions. We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.
Evaluation Research
Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions.
Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome. Evaluation Research is not conventional QA and it is not an internal approval service. It is an independent research function that works closely with the teams building Aaru's systems while preserving the ability to reach and communicate inconvenient conclusions.
role
As Evaluation Research Manager, you will lead a focused team of Evaluation Researchers and research engineers. You will translate Aaru's evaluation charter into a coherent portfolio of studies, datasets, and shared infrastructure, and you will be accountable for the quality, pace, and usefulness of the team's work. Managers at Aaru remain researchers. You will design studies, write analysis code, inspect individual failures, review statistical and measurement choices, and directly contribute to the hardest evaluations.
You will also set priorities, hire exceptional people, develop the team, provide candid feedback, and create the operating mechanisms that keep protected evidence independent while making diagnostic evidence available quickly. You will work across Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, Deployment, and company leadership. The job requires both scientific independence and practical judgment: an evaluation must be rigorous enough to support a real claim and diagnostic enough to help a team improve the system.
you will do
Build, lead, and develop a high-performing team of Evaluation Researchers and research engineers.
Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, decisive experiments, and explicit decision criteria.
Set a focused evaluation agenda across population construction, predictive systems, individual agent behavior, group dynamics, and end-to-end simulations.
Decide which evaluation infrastructure should become a reusable organizational rail and which questions require a…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).