Principal Speech Data Linguist
Listed on 2026-10-02
-
IT/Tech
AI Evaluation, Data Annotation/ AI Labeling
Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.
Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.
Scope of the Role:
As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years.
This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself.
What You’ll Own:
- You will own the linguistic foundation of Innodata's segmentation and transcription work across languages, domains, and use cases. Concretely, you will:
- Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, time stamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions.
- Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale.
- Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale.
- Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can't do rather than on what they already can.
- Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions.
- Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models.
- Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).