AI Evals Engineer — Evaluation Datasets & Ground Truth
Listed on 2026-09-09
-
IT/Tech
Data Analyst, Data Engineering, AI Evaluation, Data Annotation/ AI Labeling
Prophetic Software is not able to sponsor employment visas now or in the future. Candidates must be authorized to work in the United States without current or future sponsorship to be considered for this role.
About Prophetic:
Real estate development is a multi-billion-dollar industry that has run on fragmented data, manual processes, and gut instinct for decades. Prophetic is changing that. We're building the AI-native platform that enables homebuilders, developers, and investors to find, analyze, and act on land opportunities from a single system — powered by proprietary technologies that process billions of data points across all 50 states.
We are the market leader in our space, and our customers don't just use the product — they love it. We're not making teams more efficient. We're changing how they operate.
Prophetic is scaling fast — demand is outpacing our ability to hire, and we're just getting started. This is a once-in-a-lifetime team: sharp, low ego, deeply collaborative, and obsessed with building the best product in the industry. Come disrupt an industry with us.
Why this role existsWe iterate on pipeline components constantly: prompts, models, classifier thresholds, retrieval logic, harness design. We are hiring a full-time engineer to produce, maintain, and defend the evaluation datasets that let every meaningful component in our stack be measured. You will own ground truth at Prophetic.
This is not an evals-infrastructure role and it is not a labeling-operations role, although you’ll touch both. Your deliverable is trusted data
: for a given module, a versioned set of inputs and expected outputs, plus a written definition of what “correct” means and how confident we should be in the labels.
Decide what needs to be measured, and how
Read our pipelines and system architecture, sit with product and engineering, and decompose each system into evaluable modules with explicit input → expected-output contracts.
For each module, define what “correct” means in writing: rubrics, label schemas, edge-case policies, and the tolerances that matter to the business.
Prioritize. We have more modules than you can cover in year one; you’ll decide where a validation set unblocks the most iteration, with minimal direction from principal engineers.
Source the data by whatever means is cheapest and most trustworthy for that module
Production sampling: pull stratified, de-identified samples from real traffic so eval sets reflect what the system actually sees, including the long tail.
Human labeling: scope and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label quality, not just label throughput.
Synthetic / oracle-generated ground truth: where a task is tractable for a frontier model given enough compute (long context, multi-pass, tool use, self-consistency) but too expensive to run that way in production, design the oracle harness that produces labels for the cheap production path to be measured against. Then verify the oracle: calibrate its output against a human-labeled sample before anyone trusts it.
Programmatic and adversarial construction: heuristic labels, templated edge cases, backtests built from past incidents (“what test would have caught this?”).
Make the data trustworthy over time
Version every dataset; track lineage, splits, and which model/prompt versions have seen which examples.
Guard against contamination and leakage (eval examples drifting into few-shot prompts, the oracle model also being the production model, etc.).
Slice by customer segment, input type, and difficulty so a headline number can’t hide a regression.
Refresh sets as the product and traffic change; retire stale examples.
Close the loop with engineering
Calibrate automated graders (LLM-as-judge, similarity metrics, exact-match) against your human gold sets, and be the person who says when an automated judge is good enough to gate on.
Report metrics correctly: precision/recall/F1, confusion matrices, calibration, confidence intervals, sample sizes needed to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).