×
Register Here to Apply for Jobs or Post Jobs. X

Principal Data Engineer

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Sanas
Full Time position
Listed on 2026-06-23
Job specializations:
  • IT/Tech
    Data Engineering
Salary/Wage Range or Industry Benchmark: 125000 - 150000 USD Yearly USD 125000.00 150000.00 YEAR
Job Description & How to Apply Below

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real‑time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross‑language communication, and more.

Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language.

Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design‑minded product engineering culture to build and ship cutting‑edge AI models and experiences — entirely in‑house.

Sanas is a 180‑strong team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies.

With Sanas, you’re not just adopting a product; you’re investing in the future of communication.

If you’re looking to have a significant role in road mapping and driving technical directions, to deploy challenging and big ideas without much overhead or slowness, and to leave your mark on an ambitious, generational mission to change how the world thinks about speech + AI, then Sanas is a well‑suited place for you.

About the Role

Our models are only as good as the data that trains them. As a Staff Data Engineer, you’ll own the infrastructure that takes raw audio — millions of hours across accents, languages, noise conditions, and recording environments — and turns it into clean, reproducible, training‑ready data ’ll work directly with AI research scientists and ML engineers to design systems that move fast without breaking the data quality guarantees our models depend on.

Job Description Data pipeline & lakehouse architecture
  • Design and implement large‑scale data pipelines that ingest, transform, validate, and serve high‑quality audio and metadata for AI model training, evaluation, and product telemetry.
  • Own the lakehouse architecture — table format choices (Iceberg vs. Delta Lake), partitioning strategies, metadata management, and schema evolution — with a bias toward reproducibility and auditability.
  • Build and maintain batch and streaming pipelines using Spark, Flink, and orchestration tooling (Airflow or Dagster), with a clear‑eyed view of when each is the right tool.
  • Extend and maintain feature store infrastructure to serve low‑latency, versioned features for both training and real‑time inference.
  • Develop and maintain pipelines purpose‑built for the unique challenges of audio data: large file volumes, time‑series feature extraction, speaker and language metadata, and annotation versioning.
  • Build tooling that supports the full audio data lifecycle — from raw ingestion and quality filtering through augmentation, segmentation, and training split generation — with reproducibility guarantees at every stage.
  • Partner with ML engineers and research scientists to design data schemas, sampling strategies, and evaluation datasets that accurately reflect production conditions.
  • Own data pipelines that feed human‑in‑the‑loop annotation workflows — ensuring clean round‑trips between raw data, labeling platforms, and training‑ready outputs.
Platform reliability & governance
  • Instrument pipelines with observability, data quality checks, lineage tracking, and alerting — so failures surface fast and root causes are traceable.
  • Drive build vs. buy decisions for data quality, observability, and cataloging tooling with a clear framework grounded in Sanas's scale and roadmap.
  • Own disaster recovery design for critical data assets — training datasets, evaluation benchmarks, and model checkpoints.
Technical leadership
  • Set…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary