×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineer, Data Engineering

Job in Foster City, San Mateo County, California, 94420, USA
Listing for: WorkGenius Group
Full Time position
Listed on 2026-07-20
Job specializations:
  • Software Development
    Data Engineering
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Onsite in Foster City, CA | 5 days in office

The Autonomy Behavior ML Data Optimization team is looking for a Software Engineer with strong data processing and pipeline engineering skills to build, scale, and optimize Scenario Scout – the scenario discovery platform at our Client that empowers teams across the autonomous vehicle development lifecycle. Scenario Scout transforms large-scale driving data into actionable insights by enabling fast, semantic similarity search over millions of driving scenario embeddings.

Our Client is building a future for Riders, not drivers. At the heart of this vision, the Autonomy Behavior ML Data Optimization team develops models that forecast the behavior of all road agents around our Client's robotaxi. Scenario Scout is the useful data tool that enables teams to discover, curate, and validate driving scenarios – from finding diverse training data for ML models to investigating safety‑critical edge cases and enriching Focus Area (FA) datasets.

The core of this role is data pipeline engineering: designing and operating Airflow DAGs that orchestrate end‑to‑end dataset creation and refresh pipelines, building Ray‑based distributed processing jobs for k‑means clustering, embedding cache generation, FAISS index building, and dataset statistics computation over large Parquet datasets stored in S3. The engineer will also work on the Python/FastAPI backend powering high‑throughput embedding search and bulk operations, and contribute to the React.js

frontend for UI‑driven pipeline management, admin panel features, and search experience improvements. The ideal candidate is someone who thrives in data‑intensive environments and is comfortable working across the stack when needed, with particular depth in data processing and pipeline orchestration.

If you are excited about building data pipelines and processing systems that directly accelerate safe autonomous driving, enjoy working with large‑scale distributed data infrastructure, and want to make a measurable impact on how a world‑class AV company discovers and understands driving scenarios, this role is for you.

Responsibilities:
  • Build and maintain the Python/FastAPI backend powering high‑throughput embedding search, bulk execution APIs, and dataset management endpoints.
  • Develop, maintain, and enhance dashboards and data visualization tools for data introspection and ad‑hoc reporting.
  • Implement observability and monitoring (pipeline health, search hit rate, result download rate, system health metrics) to ensure platform reliability and data‑driven iteration.
  • Design, build, and maintain Airflow DAGs that orchestrate end‑to‑end dataset creation and refresh pipelines, including embedding generation, distributed k‑means clustering, FAISS index building, embedding cache construction, and dataset statistics computation.
  • Optimize data processing performance and pipeline reliability, including embedding generation pipeline improvements, cache building optimizations, and index construction tuning.
  • Enhance and operate scalable data processing pipelines for distributed k‑means clustering, embedding cache building, FAISS index generation, and dataset statistics computation on large Parquet datasets stored in S3.
  • Own the Scenario Scout dataset lifecycle: full dataset creation, incremental refresh, bulk clustering assignment, dataset registry management, and embedding onboarding for new embedding types (e.g., non‑QTP embeddings).
  • Collaborate with ML researchers on integrating new embedding types, improving embedding quality, and exploring LLM‑powered natural language querying capabilities.
  • Contribute to the React.js frontend for admin panel features (dataset creation forms, pipeline monitoring, metadata field configuration), search experience improvements, and UI‑driven pipeline management.
  • Deliver UI/UX improvements informed by user feedback, including visualization enhancements, visualization toggles, query result deduplication, shareable search result links, search interruption controls, and result grouping for overlapping scenarios.
  • Work closely with cross‑functional stakeholders (Prediction, Data Optimization, Safety, QA) to understand…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary