RLHF Quality Reviewer and Model Evaluator Remote English C1, French C1 CAD
Canada
Listed on 2026-09-24
-
Science
AI Evaluation, Data Annotation/ AI Labeling
Reinforcement learning from human feedback is only as good as the judgement behind it, and at scale that judgement has to be audited. You will be the audit. You will review the work of annotators, adjudicate disagreements, and identify whether a quality failure originates in an individual, in the training, or in the guideline itself.
Fully remote across Canada, minimum twenty hours weekly.
What you will doReview annotator preference judgements, ratings and written responses for correctness, guideline consistency and freedom from bias. Adjudicate disagreements between annotators with a documented rationale that a client can audit. Sample and audit batches, reporting quality metrics against programme thresholds. Distinguish systematic error from individual error and say which, because the remedies are entirely different. Give annotators feedback that measurably improves their next batch.
Contribute to guideline refinement where the current version produces inconsistent outcomes. Support red team and safety evaluation programmes, which can involve reviewing content designed to elicit harmful model output. That work is optional, clearly flagged at assignment, and supported.
English at C1, with rigorous written reasoning. At least two years in annotation, linguistic quality, editorial work, translation review, teaching, research or a comparable judgement-based discipline. Ability to reason carefully about ambiguous cases and to write down the reasoning. Sound judgement on safety, bias and factual accuracy. Total reliability on deadlines, since a delayed audit blocks a training run.
Nice to havePrior RLHF or model evaluation experience. Postgraduate qualification in linguistics, philosophy, law, cognitive science or a scientific field. French at C1, which opens bilingual programmes at a premium rate. Subject depth in a specialist domain.
What we offerRates that rise with review tier and language scarcity. Fully remote and flexible within weekly minimums. Paid guideline training per programme. Rolling contracts with steady volume. Work that directly shapes model behaviour, and a clear route into programme leadership.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).