×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist, Video & Multimodal

Job in Ridgefield Park, Bergen County, New Jersey, 07660, USA
Listing for: Innodata
Full Time position
Listed on 2026-10-01
Job specializations:
  • Science
    AI Evaluation, Data Annotation/ AI Labeling
Salary/Wage Range or Industry Benchmark: 160000 - 185000 USD Yearly USD 160000.00 185000.00 YEAR
Job Description & How to Apply Below

Innodata(Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.

Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted  provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope

of the Role

Video is where multimodal models are weakest and hardest to grade. Temporal reasoning, long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks — and the evaluations for them are still immature. Closing that gap is gated as much by how we design data and evaluation as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it.

What

You’ll Own

You will define how Innodata designs, structures, and evaluates video data for video and multimodal models, and you will validate those choices experimentally. Concretely, you will:

  • Translate the requirements of video and multimodal models — video understanding, temporal and event localization, action recognition, long-form video, video-language models, video generation, cross-modal reasoning, and multimodal retrieval and grounding — into concrete data specifications: modalities, annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology for video understanding — temporal grounding accuracy, long-context and long-horizon reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluation — clear about when model-based scoring is trustworthy and when a human is needed.
  • Build evaluation methodology for video generation — fidelity, temporal coherence, and physical plausibility, including generative video used as a world model — the regime where automatic metrics are weakest and human judgment matters most.
  • Decide how existing and incoming video should be structured, enriched, and sampled to extract the most model value from it, including from messy, domain-specific footage.
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement.
  • Design adversarial and stumping evaluations that surface where video and multimodal systems fail, and turn those failures into better data.
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with.
  • Work with annotation teams, subject-matter experts, and the synthetic-data pipeline to turn specifications into operational collection and labeling plans.
You'll Thrive in This Role If You Have
  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end.
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Trained and evaluated video or multimodal models yourself, with strong PyTorch fundamentals.
  • Fluency in the formats and tooling video work runs on: ffmpeg and decord pipelines, temporal and COCO-style annotation, Web Dataset, Parquet and Arrow, and Hugging Face datasets.
  • Experience…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary