Data Platform Engineer
El Segundo, Los Angeles County, California, 90245, USA
Listed on 2026-09-23
-
IT/Tech
Data Engineering
Your role is to build the infrastructure that collects, stores and models everything the company knows about its customers, its machines and itself, and turns it into decisions.
That infrastructure is Cyberdeck: we have rebuilt in our own image every operational software product a company like ours would buy, and run it on our own hardware. Owning the software keeps the data in one estate under one identity model for people and agents.
You will own the stores and the pipelines that fill them, quality and provenance as properties of the pipeline, the entity model beneath it, the time-series and geospatial work the company asks of that data, and the serving layer people and agents read through.
We are looking for data engineers with experience and interest in building and operating production data platforms, information modeling and entity resolution, provenance and data quality, time-series pipelines, and geospatial analysis.
Every role works with our AI systems daily. What can be deterministic, must be. You need no AI background; we prefer people without one. We hire for your knowledge and experience in the field first, so you can steer the ship; the tooling is a learning curve we expect you to take on.
We work on-site in El Segundo. This is hard work, but you will be rewarded with equity in a company we believe will become one of the world's most valuable.
What You'll Do- The data estate: stores, schemas and indexes, migrations, query performance, retention and isolation rules
- Ingestion pipelines: web enrichment, customer orders, machine telemetry, accounting feeds, idempotent, replayable and backfillable
- Data quality: contracts and tests at the boundary, anomaly detection, freshness ranking, stewardship workflows
- The canonical entity model: one identity per person, company, design, part, order and machine
- Entity resolution: deterministic keys, blocking and scoring, a review queue, and reversible merges
- Provenance: source, confidence and history on derived records, and corrections that survive a rerun
- Time-series and events: a taxonomy for the floor, downsampling, retention tiers, and reliability metrics
- Geospatial and serving: geocoding, drive-time catchments, facility siting near customers, an analytics layer
- A degree in computer science, data engineering, information systems or a related field, or equivalent experience
- Hands-on production data pipelines you were responsible for: idempotency, replay, backfills and schema evolution
- Information modeling and entity resolution someone else had to live with: canonical identifiers, matching, reversible merges
- Fluent with document and relational stores, SQL and Python, and able to explain a slow query
- The judgment to take a wrong number back through the pipeline instead of correcting the output
- Ready to build the first version of a system and write it for the next engineer
- Bonus: knowledge graphs or taxonomy design, geospatial work such as isochrones and catchments, or high-rate time-series
$100k to $250k base salary, plus equity
- Equity participation in a high-growth startup
- Comprehensive health, dental, and vision insurance
- 401(k) with company matching
- Bonus for living within 5 miles of our El Segundo facility
- On-site work with rare work-from-home exceptions
- Merit-based organization where contribution drives reward
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).