Volunteer AI Data Annotator; Benchmarks & Communication Rubrics AI Apps Human Communication
Listed on 2026-10-11
-
Business
AI Evaluation
Location: New York
Crossing Party Lines (CPL) provides opportunities, training, and support for civil, respectful discourse between Americans with dissimilar ideologies. Our primary offering is the CPL meetup regularly scheduled, facilitated meetings to which all are welcome. At our meetups, attendees discuss current hot topics while facilitators focus the meetings to teach and model skills and create a safe space for members to
- Be heard by and listened to by people with different political views.
- Express views without negative repercussions.
- Learn from one another through active listening.
- Develop curiosity about perspectives that may be new and unfamiliar to them.
- Find connection with members who espouse different views.
- Develop skills required for civil, respectful discourse.
In 2016, two groups named "Crossing Party Lines" started independently of one another on opposite ends of the country, only a few months apart. They joined forces to bring civil, respectful conversation to the rest of the country.
Review, label, and score conversational AI data to help establish gold-standard benchmarks for tools designed to reduce political polarization and foster constructive dialogue.
Crossing Party Lines is seeking a thoughtful, detail-focused AI Data Annotator to join our Apps and Agents team. Our conversational agents are built to handle emotionally sensitive, politically nuanced interactions. Standard automated evaluators often miss subtle conversational missteps, which makes human ground-truth data indispensable. In this role, you will analyze simulated dialogues, evaluate AI responses against qualitative communication rubrics, and generate labeled benchmark datasets used to train, evaluate, and align our models.
AboutThe Apps And Agents Team
Our team integrates generative AI with Crossing Party Lines’ proprietary frameworks developed through nearly a decade of real-world dialogue facilitation. As we prepare an application to Y Combinator and expand our suite of digital role-play and coaching tools, high-quality human data is essential to ensure our agents remain objective, emotionally safe, and genuinely depolarizing.
What You’ll Do- Review multi-turn conversational transcripts between users and AI agents across diverse, high-friction social and political topics.
- Annotate and score AI outputs using structured rubrics covering active listening, perspective-taking, neutrality, de-escalation, tone calibration, and defensiveness triggers.
- Identify and tag failure modes such as subtle ideological bias, sycophancy, tone-policing, preachy phrasing, and hallucinations.
- Write and curate high-quality "golden responses" (ideal model outputs) to serve as few-shot exemplars and ground-truth references for evaluation pipelines.
- Provide qualitative feedback to AI engineers, prompt designers, and subject-matter experts on recurring model behaviors and annotation edge cases.
- Participate in inter-annotator calibration sessions to maintain high data consistency and label reliability across the team.
- Have strong reading comprehension, analytical judgment, and close attention to language nuance and subtext.
- Bring an objective, open-minded approach capable of evaluating dialogue across a wide spectrum of political and social viewpoints without personal bias.
- Have an interest in linguistics, psychology, communication studies, political science, philosophy, or AI safety.
- Enjoy repetitive, analytical tasks that require high precision, consistency, and adherence to labeling guidelines.
- Prior experience with data annotation platforms (Label Studio, Prodigy, or custom spreadsheets) is a plus, but not required.
- Practical experience in human-in-the-loop (HITL) AI development and LLM benchmarking methodologies.
- A direct role in setting the safety and alignment standards for civic technology applications.
- Collaboration with AI engineers, data specialists, and conflict-resolution experts.
- Professional references recognizing your attention to detail, analytical rigor, and contributions to production AI evaluation datasets.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).