×
Register Here to Apply for Jobs or Post Jobs. X

Staff Engineer, Autonomous Driving Data Platform & Curation

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: CARIAD, Inc.
Full Time position
Listed on 2026-09-25
Job specializations:
  • Science
    Data Annotation/ AI Labeling
Salary/Wage Range or Industry Benchmark: 162000 - 234000 USD Yearly USD 162000.00 234000.00 YEAR
Job Description & How to Apply Below

Role Summary

The Staff Engineer, Autonomous Driving Data Platform & Curation is a hands‑on staff‑level individual contributor who owns the path from raw multimodal vehicle data to reliable, versioned, and model-ready datasets. The ideal candidate combines production data or ML systems experience with practical knowledge of autonomous‑driving data and can lead ingestion, schema, storage, query, curation, quality, and delivery. This engineer designs and operates datasets containing more than 10 million records or samples, evaluates formats such as Parquet and Lance, and enables efficient filtering, slicing, random access, and sequential retrieval.

The role also advances data‑quality monitoring, statistical and out‑of‑distribution detection, rule‑based and model‑based tagging, model‑in‑the‑loop and human‑in‑the‑loop labeling, and future data preparation for imitation learning and reinforcement learning.

Role Responsibilities
  • Design canonical representations for drives, scenarios, clips, frames, trajectories, sensor references, vehicle state, map context, labels, predictions, and dataset manifests.
  • Build and maintain validated, versioned datasets containing more than 10 million records or samples on local or cloud object storage.
  • Evaluate Parquet, Apache Arrow, Lance, and related technologies; own partitioning, indexing, file sizing, compaction, schema evolution, lineage, and reproducibility.
  • Optimize filtering, projection, joins, scenario slicing, random sampling, shuffling, sequential retrieval, and model data‑loading performance.
  • Measure and improve ingestion throughput, query latency, training throughput, storage utilization, reliability, and cost per usable sample.
  • Work effectively with synchronized camera and other sensor data, ego state, localization, calibration, coordinate frames, map context, control actions, clips, and trajectories.
  • Translate perception, planning, VLA, and evaluation needs into schemas, searchable attributes, scenario definitions, sampling strategies, and reproducible dataset splits.
  • Build reliable batch or distributed pipelines for ingestion, transformation, enrichment, validation, cataloging, and publication, including retries, backfills, idempotency, and observability.
  • Enable self‑service discovery and composition for lane keeping, lane changes, long‑tail scenarios, hard examples, balanced datasets, and leakage‑resistant train, validation, and test splits.
  • Define automated quality gates, dashboards, and alerts for completeness, validity, freshness, duplication, synchronization, calibration, corruption, label integrity, coverage, balance, and cost.
  • Apply statistics, sampling, and distribution comparisons to detect drift, anomalies, underrepresented conditions, and out‑of distribution data.
  • Use deterministic checks, heuristic rules, geometry and metadata queries, embeddings, VLMs, and learned models to assess quality and create searchable scenario tags.
  • Quarantine suspicious data and lead root‑cause analysis across collection, synchronization, schema, transformation, storage, annotation, sampling, and model‑consumer failures.
  • Design model‑in‑the‑loop and human‑in‑the‑loop labeling workflows with confidence thresholds, review routing, audit sampling, disagreement handling, and label provenance.
  • Close the loop from model failures and edge cases through selection, annotation, quality review, dataset publication, training, and evaluation.
  • Evaluate auto‑labeled and synthetic data using label quality, coverage, distributional impact, and downstream model performance.
  • Prepare replayable trajectory data for imitation learning and offline RL, including observations, actions, timestamps, policy versions, interventions, rewards, termination conditions, and alignment checks.
  • Serve as the…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary