×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist, Speech & Audio

Job in Ridgefield Park, Bergen County, New Jersey, 07660, USA
Listing for: Innodata
Full Time position
Listed on 2026-10-02
Job specializations:
  • Science
    Data Annotation/ AI Labeling, AI Evaluation
Salary/Wage Range or Industry Benchmark: 160000 - 185000 USD Yearly USD 160000.00 185000.00 YEAR
Job Description & How to Apply Below

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.

Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted  provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope

of the Role

Where models actually differ now is robustness across accents, noise, and code-switching; speaker diarization; the naturalness of generated speech; and latency under streaming. Measuring those honestly, and building the data that trains for them, is gated as much by data and evaluation design as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing speech and audio models, and we are hiring a Research Scientist to own the science behind it.

You will partner directly with the customers and frontier labs building ASR, text-to-speech, speech-to-speech and conversational voice, diarization, and audio-language models, as interested in the data behind them as in the models themselves. Your work is judgment: which conditions and languages a benchmark must cover to be honest, what a transcription convention should be for a given objective, and when an automated metric can be trusted versus when a human ear is required.

You will also partner closely with our transcription and linguistics lead, whose standards directly shape what the models learn.

What You’ll Own
  • You will define how Innodata designs, structures, and evaluates audio data for speech and audio models, and you will validate those choices experimentally. Concretely, you will:
  • Translate the requirements of speech and audio models — ASR, text-to-speech and speech generation, speech-to-speech and conversational voice, speaker diarization and verification, audio-language models, and streaming systems — into concrete data specifications: modalities, transcription and annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology that goes past word error rate — semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency — and know when automated metrics hold and when they don't.
  • Decide how existing and incoming audio should be structured, enriched, and sampled for coverage that fits the model objective — across languages, accents, and acoustic conditions (studio, real-world, telephonic), speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech.
  • Partner with the transcription and linguistics lead to turn model objectives into transcription specifications, and to quantify how transcription conventions and quality move ASR and speech-model results.
  • Partner with the audio solutions and engineering team so the audio we collect is built for the model objective: you specify what good data and evaluation require, and they scope programs with customers and capture audio to spec.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations — noisy, accented, and adversarial audio — that surface where speech systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary