Traditional Chinese AI Response Rater - Multi-Domain QA (Remote
Phoenix, Maricopa County, Arizona, 85003, USA
Listed on 2026-08-12
-
IT/Tech
AI Evaluation, Chinese Speaking, Data Annotation/ AI Labeling
Traditional Chinese AI Response Rater
- Multi-Domain QA (Remote) is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured rationale that the modeling team can act on in the next training cycle.
Multilingual model quality depends on reviewers who can read Chinese the way a native speaker does. Aura One uses these reviewers to set the bar for tone, formality, and cultural register before a model ever ships to Chinese users.
Responsibilities- Evaluate Chinese model responses for fluency, accuracy, register, and cultural fit for Traditional Chinese AI Response Rater
- Multi-Domain QA (Remote) assignments. - Compare paired chinese evaluation outputs and pick the stronger response with a written rationale.
- Flag hallucinations, code-switching errors, and formality mismatches with task-specific severity tags.
- Capture native-speaker edits that demonstrate the correct phrasing alongside each issue.
- Calibrate against the Aura One Chinese rubric in weekly reviewer-quality cycles.
- Surface ambiguous prompts back to the program team so the rubric can be sharpened.
- Maintain reviewer notes that document tricky idioms, regionalisms, and emerging usage.
- Native or near-native Chinese fluency with strong written-English ability for Traditional Chinese AI Response Rater
- Multi-Domain QA (Remote) work. - Demonstrable experience editing, translating, or evaluating Chinese content.
- Comfort applying multi-page rubrics consistently across long evaluation batches.
- Clear written reasoning that names the issue and the standard being applied.
- Reliable async availability for at least 10 hours per week.
- Prior model-evaluation, annotation, or human-rater experience is a plus.
- Score a pair of Chinese responses to a customer-support prompt and pick the better one with a 1-2 sentence rationale.
- Flag formality mismatches in chinese evaluation outputs (T/V usage, honorifics, regional register).
- Rewrite a model's Chinese response so it reads naturally and tag the original error category.
- Audit a 50-row evaluation batch for rubric consistency and surface drift back to the program lead.
- Background in Chinese linguistics, journalism, or translation studies.
- Experience with TMS tools, MQM error taxonomies, or other structured QA frameworks.
- Familiarity with LLM evaluation rubrics and inter-rater agreement workflows.
- Multilingual evaluation
- Native-speaker review
- Rubric application
- LLM output evaluation
- Chinese evaluation
Remote — US-eligible. Remote
· Independent specialist contractor.
Employment type:
CONTRACTOR. Applicants must be authorized to work from US.
$40–$49 / hr
Native or near-native Chinese fluency with strong written-English ability for Traditional Chinese AI Response Rater - Multi-Domain QA (Remote) work. Demonstrable experience editing, translating, or evaluating Chinese content. Comfort applying multi-page rubrics consistently across long evaluation batches. Clear written reasoning that names the issue and the standard being applied. Reliable async availability for at least 10 hours per week. Prior model-evaluation, annotation, or human-rater experience is a plus.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).