×
Register Here to Apply for Jobs or Post Jobs. X

Principal Engineer, Model Development Platform

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Icehouseventures
Full Time position
Listed on 2026-07-06
Job specializations:
  • Software Development
    Software Architect, Cloud Engineer - Software, DevOps, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 295500 - 335300 USD Yearly USD 295500.00 335300.00 YEAR
Job Description & How to Apply Below

About us

Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and safety of automated driving systems.

Our vision is to create autonomy that propels the world forward. Our intelligent, mapless, and hardware-agnostic AI products are designed for automakers, accelerating the transition from assisted to automated driving.

In our fast-paced environment big problems ignite us—we embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.

At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive work environment; we back each other to deliver impact.

Make Wayve the experience that defines your career!

As Principal Engineer for the Model Development Platform, you'll own the end-to-end architecture behind Wayve's AI model lifecycle, from data ingestion and training to experiment scheduling and on-road testing. Working at the intersection of AI research, large-scale distributed systems, and robotic operations, you'll keep the platform reliable, scalable, and coherent so our researchers and engineers can iterate fast and deploy autonomous driving models safely.

Partnering with the Head of Model Dev Platform, you'll set and execute the technical vision, aligning infrastructure and tooling with company goals. You'll lead by example, going deep across web applications, distributed compute, ML Ops, data pipelines, and optimization algorithms, and through architecture and mentorship you'll enable teams to build platform capabilities that measurably accelerate model development and fleet learning.

What you'll own

  • System architecture & reliability - Design and evolve the platform's overall architecture for reliability, observability, and scalability. Set performance, latency, and availability targets, and drive the engineering standards to meet them.
  • Cross-domain technical leadership - Unify the platform across disciplines, from front-end UIs and distributed training to Spark data pipelines and optimization-based experiment scheduling, ensuring systems interoperate cleanly.
  • Hands-on problem solving - Dive into the hardest challenges across subteams, lead architectural reviews, and propose pragmatic solutions that balance innovation with operational simplicity.
  • Experimentation & scheduling systems - Build systems that optimize how models are tested in simulation and on-road, using techniques like linear programming and heuristic optimization to balance hardware, safety, and research priorities while improving throughput and turnaround.
  • Data & compute infrastructure - Architect pipelines that ingest, transform, and enrich petabytes of fleet sensor data, and drive efficient compute use across GPU, CPU, cloud, and edge for both prototyping and large-scale training.
  • Strategic collaboration - Partner with Product, Research, and Operations to align architecture with user needs and co-own the platform's long-term roadmap.
About You

Essential

  • Technical Leadership at Scale – 10+ years of experience designing and building large-scale distributed systems, ML/AI infrastructure, full stack web application, or developer platforms, including at least 3 years as a staff or principal-level engineer.
  • Architectural Depth & Breadth – Proven ability to design systems spanning web platforms, ML pipelines, and large-scale compute orchestration (e.g., Spark, Ray,
    Kubernetes
    , Airflow, MLflow).
  • Reliability and performance – Experience driving platform reliability improvements, defining SLAs/SLOs, and building self-healing and observable systems that operate at “four nines” availability or better.
  • Hands-On Systems Design – Deep understanding of distributed computing, workflow orchestration, data modeling, and API design, with the ability to write and review production-quality code.
  • Collaborative Influence – Excellent communication and cross-functional collaboration
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary