×
Register Here to Apply for Jobs or Post Jobs. X

Evaluation Researcher

Job in New York, New York County, New York, 10261, USA
Listing for: Aaru
Full Time position
Listed on 2026-08-28
Job specializations:
  • Research/Development
    Research Scientist, Data Scientist, Research Analyst, AI Evaluation
Salary/Wage Range or Industry Benchmark: 120000 - 190000 USD Yearly USD 120000.00 190000.00 YEAR
Job Description & How to Apply Below
Location: New York

About Aaru

Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes. Building a useful simulation requires more than generating plausible text.

Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions. We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.

About

Evaluation Research

Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions.

Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome. Evaluation Research is not conventional QA and it is not benchmark administration. It is an independent research function. The work requires understanding the systems deeply, collaborating closely with their builders, and remaining willing to conclude that an attractive method did not improve what matters.

The

role

As an Evaluation Researcher, you will own difficult measurement problems at the boundary of machine learning, statistics, behavioral science, and product decision-making. You will define constructs, assemble or create evaluation data, design studies, write analysis and evaluation code, inspect individual failures, quantify uncertainty, and communicate what the evidence does and does not support. Some projects will build reusable rails used across the research organization.

Others will be focused studies intended to resolve one important uncertainty. In both cases, the goal is the same: create an evaluation that is valid enough to trust, diagnostic enough to guide improvement, and clear enough to inform a real decision. You will work closely with Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, and Deployment while protecting the independence and integrity of final measurements.

What

you will do

Own an important evaluation or measurement area across population construction, predictive systems, individual agent behavior, group dynamics, or end-to-end simulations. Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs and explicit decision criteria. Design studies using historical backtests, temporal holdouts, prospective outcomes, observational records, controlled experiments, expert judgment, or mixed methods as the problem requires. Build tests of individual-profile quality, including internal coherence, contradictions, impossible combinations, unsupported specificity, stability, and whether a profile induces behavior consistent with the represented person.

Build tests of population quality, including marginal and joint distributions, conditional relationships, coverage of rare but plausible profiles, subgroup fidelity, and sensitivity to sampling choices. Evaluate predictions using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the decision cost of different errors. Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, surveys, and longitudinal decisions.

Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another. Build end-to-end studies that determine whether a component improvement actually changes the quality of the conclusion a customer receives. Establish strong baselines and compare agent-based simulation with direct forecasting, conventional statistical models, simpler segment-level methods, and human or market benchmarks where appropriate.

Find failures hidden by aggregate metrics, especially failures concentrated in important subgroups, rare cases,…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary