Biomedical Data Annotator | Remote
Northern, Floyd County, Kentucky, USA
Listed on 2026-09-22
-
Research/Development
AI Evaluation, Clinical Research, Research Scientist
At a Glance
We need human curators to develop a range of biomedical question and answer datasets for AI research. Learn how trustworthy AI systems are built and contribute your scientific expertise to the National Library of Medicine (NLM) at the National Institutes of Health (NIH).
About Mind Moves
Mind Moves is a women-owned Washington, D.C.
-based firm that helps government and business partners navigate digital transformation using "human-in-the-loop" AI. Our expert team has delivered responsibly developed AI products that drive millions in impact across agencies like the National Institutes of Health (NIH).
About the Role
Mind Moves is seeking 20+ independent Data Annotators to develop a benchmark evaluation dataset for AI systems operating over the biomedical literature. As these systems become integrated into research and clinical workflows, rigorous assessment of their safety, factual accuracy, and scientific reliability requires expert-authored ground truth.
SMEs will independently create authentic biomedical information needs and reference responses grounded exclusively in evidence from scientific literature resources Pub Med and Pub Med Central (PMC), following standardized annotation guidelines. The resulting benchmark will provide a transparent, reproducible resource for evaluating AI systems’ ability to retrieve, synthesize, and accurately communicate biomedical evidence.
Contract Engagement
This is a project-based, hourly contract engagement supporting a defined research project in AI-powered health. This is a short-term contract assignment and is not intended to lead to full-time regular employment.
Engagement Details
Project Duration: September–November 2026
Expected Commitment: Approximately 10-20 hours per week, with a flexible remote schedule.
Compensation: $30 per hour.
Contract
Location:
Remote, based in the US.
Expert Profiles Sought
We are contracting against two distinct expert profiles detailed here:
Profile 1— Academic and Bench Researchers
PhD candidates (post-candidacy, actively engaged in dissertation research) and early-stage postdoctoral researchers with current, deep command of their literature. Sub-domains may include, but not limited to, molecular genetics, structural biology, immunology, oncology, multi-omics, and computational biology.
Profile 2— Research-Savvy Clinicians and Healthcare Professionals
Medical, pharmacy, nursing, and allied health students, along with early-career clinicians and healthcare professionals, who regularly interpret primary biomedical literature for research or evidence-based clinical practice. Specialties sought include, but not limited to, cardiology, oncology, neurology, infectious disease, pharmacology, and subgroup-specific therapeutics.
Human Origination Requirement
Because this benchmark is intended to measure AI performance against expert human reasoning, every deliverable must be entirely human authored. Work product that was drafted, rephrased, summarized, snippet-extracted, or brainstormed with any large language model — public, commercial, private, or locally hosted (ChatGPT, Claude, Gemini, Copilot, Deep Seek, Llama, Mistral, or any comparable conversational model or API endpoint) — does not conform to the specification and cannot be accepted.
Each submitted block is accompanied by a short certification of human origination.
Search engines, corpus keywords and retrieval filters, reference managers such as End Note and Zotero, and consultation with medical librarians and colleagues are all consistent with the specification. Note that literature searches must be run as decomposed conceptual queries — gene aliases, drug targets, mechanistic keyword pairs — rather than by submitting a topic's full question string verbatim, as verbatim submission risks contaminating the benchmark.
Interview ProcessApplications are accepted on a rolling basis, though we look to fill these roles as soon as possible. We expect the process to include a short hiring exercise followed by 1 20-30 min interview for final candidates.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).