×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist – Speech and Audio Understanding; Models & Multimodal Systems

Job in Bellevue, King County, Washington, 98009, USA
Listing for: Lightspeed Studios
Full Time position
Listed on 2026-09-01
Job specializations:
  • IT/Tech
    AI Business & Operations, Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 122500 - 229700 USD Yearly USD 122500.00 229700.00 YEAR
Job Description & How to Apply Below
Position: Research Scientist – Speech and Audio Understanding (Large Models & Multimodal Systems)

What the Role Entails

We are building large‑scale, native multimodal model systems that jointly support vision, audio, and text to enable comprehensive perception and understanding of the physical world.

Job Responsibilities
  • Develop general‑purpose, end‑to‑end large speech models covering multilingual automatic speech recognition (ASR), speech translation, speech synthesis, paralinguistic understanding, and general audio understanding.
  • Advance research on speech representation learning and encoder/decoder architectures to build unified acoustic representations for multi‑task and multimodal applications.
  • Explore representation alignment and fusion mechanisms between audio/speech and other modalities in large multimodal models, enabling joint modeling with image and text.
  • Build and maintain high‑quality multimodal speech datasets, including automatic annotation and data synthesis technologies.
Who We Look For
  • Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Linguistics, or a related field; or Master’s degree with several years of relevant experience.
  • Solid understanding of speech and audio signal processing, acoustic modeling, language modeling, and large model architectures.
  • Proficient in one or more core speech system development pipelines such as ASR, TTS, or speech translation; experience with multilingual, multitask, or end‑to‑end systems is a plus.
  • Experience with large‑scale training and distributed systems is a plus.
  • Familiar with Transformer‑based architectures and their applications in speech and multimodal training/inference.
Preferred Expertise
  • Speech representation pretraining (e.g., HuBERT, Wav2

    Vec, Whisper).
  • Multimodal alignment and cross‑modal modeling (e.g., audio‑visual‑text).
  • Experience driving state‑of‑the‑art (SOTA) performance on audio understanding tasks with large models.
  • Proficient in deep learning frameworks such as PyTorch or Tensor Flow.
Location

State(s): US-Washington-Bellevue

Compensation

The expected base pay range for this position is $ – $ per year. Actual pay may vary depending on job‑related knowledge, skills, and experience.

Benefits

Employees may be eligible for a sign‑on payment, relocation package, restricted stock units, medical, dental, vision, life and disability benefits, and participation in the company’s 401(k) plan.

Vacation: up to 15 to 25 days per year (depending on tenure). Holidays: up to 13 days per year. Paid sick leave: up to 10 days per year.

Equal Employment Opportunity

We are an equal‑opportunity employer. We firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. Every employee of Tencent feels supported and inspired to achieve individual and common goals.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary