×
Register Here to Apply for Jobs or Post Jobs. X

AI Engineer, Agentic Voice; TTS

Job in Kitchener, Ontario, Canada
Listing for: Dialpad
Full Time position
Listed on 2026-09-03
Job specializations:
  • Software Development
    AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 145000 - 172500 CAD Yearly CAD 145000.00 172500.00 YEAR
Job Description & How to Apply Below
Position: AI Engineer, Agentic Voice (TTS)

About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.

Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile.

Being a Dialer

At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.

We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.

We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits:
Scrappy, Curious, Optimistic, Persistent, and Empathetic
.

Your role

As an AI Engineer on our Speech Team, you'll own the back-end implementation and linguistic optimization of the voice (TTS) layer for our next-generation AI agents. You'll work squarely within our Speech Team, a high-impact R&D and engineering group focused on speech recognition, enhancement, and synthesis; bridging core speech science and product engineering so our agents sound human, context‑aware, and trustworthy.

You’ll build the systems that render voice: integrating and optimizing TTS engines, engineering the persona and parameter machinery, and exposing voice attributes to our customer‑facing UI. You'll partner closely with our Voice Experience Designer, who authors and owns the persona standards and quality bar that you implement — you own the platform that makes their designs real, fast, and consistent at scale.

This position reports to our Senior Manager, AI Speech, is based at our Kitchener hub, and operates on a hybrid schedule.

What you’ll do
  • TTS backend implementation:
    Own the integration and optimization of multiple TTS vendor APIs behind a unified interface with failover, and lead research and prototyping of open‑source and in‑house TTS architectures.
  • Latency & pipeline engineering:
    Minimize time‑to‑first‑audio and end‑to‑end latency across the ASR → LLM → TTS pipeline while maintaining or optimizing voice quality, in partnership with ASR and Audio AI engineers.
  • Linguistic optimization:
    Apply your knowledge of phonetics and sociolinguistics to format TTS input for maximum naturalness — SSML tags, punctuation‑driven prosody, and text normalization for names, numbers, dates, and currency.
  • Persona system & parameter exposure:
    Build the persona parameterization system and architect the logic that exposes voice attributes to the product UI, implementing the house standards defined by Agent Experience Design.
  • Prompt engineering as code:
    Manage structured LLM and TTS prompt templates with versioning and a rigorous evaluation harness.
  • Conversational turn design:
    Engineer context- and state‑aware "thinking" utterances that maintain caller trust while tool calls and model steps run under the hood.
Skills you’ll bring
  • Technical foundation:
    Strong Python and hands‑on experience with deep learning frameworks (e.g. PyTorch).
  • Speech expertise: 3+ years in Speech Synthesis (TTS) or applied speech ML, including hands‑on work with frameworks like NVIDIA NeMo, ESPnet, or Coqui, and with major TTS APIs such as Eleven Labs, Rime, and Cartesia.
  • Linguistic background:
    D…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary