Research Scientist
Listed on 2026-09-13
-
Research/Development
-
IT/Tech
Build expressive speech foundation models at trillion-parameter scale, with real world impact.
This team has built their entire speech stack in-house, including proprietary LLM-based ASR and TTS, both already outperforming SOTA benchmarks. They're already powering real-time, human-to-AI conversations at scale, every day.
Now they're working at trillion-parameter scale to push toward a genuine end-to-end speech-to-speech LLM, one that understands and responds with genuine emotional intelligence and able to have natural, human-like conversations.
That means solving problems like long-context reasoning, pronunciation accuracy and maintaining consistency in noisy, real-world environments, in a domain where no model has cracked this yet.
They're hiring Senior and Staff-level Speech Scientists (Principal-level also considered) to help drive this frontier work.
What you'll do:- Build SOTA speech models from the ground up, at genuinely large scale
- Own problems end-to-end, from research through to production
- Solve hard, domain-specific speech challenges as part of the push toward speech-to-speech LLM research
- Help shape technical direction as part of a team working at the frontier of this space
- Deep, hands-on expertise in at least one of:
Speech/Audio LLMs, TTS or audio generation, or speech-to-speech - Experience shipping speech systems at real scale, not just research-stage prototypes
- Experience with pre-training speech foundation models (HuBERT, Wav2
Vec or similar) - Experience with Multimodal LLMs especially to reduce hallucinations
- Experience with Speech LLMs for low latency streaming ASR
Up to $350K-$400K base (DOE), plus substantial equity.
This is an onsite position in South Bay or Seattle.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).