Linguist III
Listed on 2026-09-07
-
IT/Tech
Data Analyst, Data Annotation/ AI Labeling, AI Evaluation, Data Scientist
Linguist (Linguistic Engineering & LLM Data Quality)
Location:
Fully Remote (US Only) Contract Duration: 12 months Pay Rate: $53 - $55/hr on W2
We are looking for a skilled Linguist III to join our team and help build, maintain, and analyze the datasets that power next-generation LLM-driven features for smart glasses and wearable technology.
In this role, you will play a critical part in building and assessing the quality of data processed by production LLM systems. You will work at the intersection of linguistics, data annotation, experimentation, and AI quality evaluation, partnering closely with engineers, research scientists, project managers, and data scientists.
You will help translate complex cross-functional requirements into high-quality datasets, develop and improve annotation workflows, analyze model and rater performance, and provide actionable insights that improve the accuracy, reliability, and overall quality of AI-powered experiences.
This is an opportunity to apply your linguistic expertise to rapidly evolving AI systems while expanding your technical capabilities in Python, SQL, statistics, data analysis, and LLM evaluation.
Responsibilities- Aggregate requests from cross-functional partners and translate them into actionable, high-quality datasets using synthetic data collections and samples from live user traffic.
- Build, curate, and maintain datasets used for LLM training, evaluation, and quality measurement.
- Develop and maintain manual and automated data-quality processes across multiple concurrent projects.
- Create and optimize LLM quality-grading queues and component-level error-attribution queues.
- Define grading rubrics, write annotation guidelines, onboard and calibrate raters, and continuously improve queue design.
- Design and conduct experiments to evaluate rater quality, guideline clarity, inter-rater reliability, and model performance.
- Analyze numerical rating results and identify meaningful trends, patterns, and quality issues.
- Conduct daily quality audits of annotated datasets, identify inconsistencies, and investigate sources of annotation and model-quality issues.
- Attribute errors to specific LLM components and dimensions, including factuality, brevity, coherence, and related quality metrics.
- Develop datasets and author guidelines rapidly in response to new and emerging LLM data-rating, creation, and annotation requirements.
- Use SQL and Python to extract data, calculate metrics, analyze agreement scores, and support reporting and automation.
- Analyze system-level and subcomponent performance metrics and translate findings into clear recommendations.
- Prepare summary reports and communicate results, milestones, risks, and recommendations to cross-functional stakeholders.
- Collaborate closely with engineering, research, product, project management, and data science teams to align on data priorities and inform product decisions.
- Operate effectively in a fast-paced environment where LLM capabilities, evaluation requirements, and priorities evolve quickly.
Dataset Building & Curation:
Receive requests from engineers and research scientists, then design and execute synthetic data collections or sample existing data to produce clean, structured, and labeled datasets for LLM training and evaluation.
Annotation Queue Management:
Stand up and maintain grading queues by defining rubrics, developing annotation guidelines, onboarding raters, and iterating on queue design as new AI capabilities emerge.
Quality Auditing & Error Attribution:
Review annotated data, identify inconsistencies, conduct inter-rater reliability checks, and trace quality issues to specific LLM components or performance dimensions.
Data Analysis & Reporting:
Use SQL and Python to retrieve and analyze metrics, calculate agreement scores, identify trends, and support dashboards and recurring reporting. Leverage AI-assisted tools where appropriate to accelerate query and scripting workflows.
Experiment Design:
Plan and execute controlled experiments focused on rater calibration, guideline effectiveness, annotation quality, and model performance. Summarize results and present recommendations…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).