Senior Data Engineer
Listed on 2026-07-24
-
Software Development
Data Engineering
The core mission of the Senior Data Engineer:
You will own and evolve our data pipelines on GCP — building new ones, hardening existing ones, improving data quality, and making clean, trustworthy data available across the organisation. You'll work end-to-end on streaming and batch pipelines, from CDC and event ingestion through transformation, serving, and the feature layer that powers our ML and AI products.
Your day-to-day will include designing ELT/ETL processes on Big Query and Click House, building real-time pipelines on Pub/Sub and Kafka with Dataflow (and where it fits, Flink/Spark), orchestrating workflows with Airflow, and ensuring data is properly cleaned, modelled, and served for analytics, ML training, and online inference. You'll partner with ML engineers on feature pipelines, monitoring data drift, and keeping models well‑fed and retrained as needed.
You'll consume and build REST APIs, integrate with third‑party SaaS sources, and treat infrastructure as code.
Location & Work Mode:
- US
- Fully Remote
You will be part of the Data, Analytics & AI team, collaborating closely with Infrastructure, Software Engineering, Product, and ML/AI engineers. We’re in the middle of a GCP‑native modernisation — migrating away from Snowflake toward Big Query, Bigtable, Pub/Sub, and Dataflow — so we’re looking for someone who is opinionated about clean architecture, allergic to over‑engineering, and comfortable owning systems end‑to‑end.
If retiring a legacy warehouse and standing up its replacement sounds like a good time, you’ll fit right in.
- Cloud & warehouse: GCP, Big Query, Bigtable, Cloud Storage
- Streaming & messaging:
Pub/Sub, Kafka - Processing:
Dataflow (Apache Beam), with Flink/Spark where appropriate - Orchestration:
Airflow (Cloud Composer) - Analytical store:
Click House - Languages:
Python, SQL - Modelling & quality: dbt, data quality gates
- Containers & CI/CD:
Docker, Kubernetes, Git Hub Actions or equivalent - Legacy being retired:
Snowflake
- Design and build batch and streaming pipelines on Dataflow, Pub/Sub, and Kafka feeding Big Query, Bigtable, and Click House
- Help drive the migration off Snowflake onto our GCP‑native stack — and retire shadow pipelines along the way
- Own the orchestration layer in Airflow, including SLAs, retries, and data quality gates
- Model data for analytics and for ML — including feature pipelines that serve both training and low‑latency online inference
- Partner with ML engineers on feature stores, drift monitoring, and retraining workflows
- Capture requirements from stakeholders and translate them into pragmatic, well‑scoped data products
- Continuously improve data quality, reliability, observability, and cost efficiency
- Identify new data sources worth acquiring and integrate them cleanly
- Strong data modelling and warehouse architecture skills (dimensional modelling, event‑driven, lakehouse patterns)
- Hands‑on experience with GCP data services — Big Query is a must;
Pub/Sub, Dataflow, Bigtable, Cloud Composer are strong pluses - Production experience with streaming pipelines on Dataflow/Beam, Flink, or Spark Structured Streaming, ingesting from Kafka and/or Pub/Sub
- Solid SQL and strong Python — you write production‑quality code, not just notebooks
- Experience with Click House or another columnar OLAP engine in production
- Workflow orchestration experience with Airflow (or Prefect/Dagster)
- Comfortable with dbt or equivalent transformation frameworks
- Experience migrating off legacy warehouses (Snowflake, Redshift, Synapse) onto cloud‑native stacks is a plus
- Working knowledge of ML in production — feature engineering, feature stores, model deployment, drift monitoring, retraining
- Docker & Kubernetes experience
- CI/CD mindset, infrastructure‑as‑code sensibility, and a bias for simple, observable systems
- Bonus: CDC tooling (Datastream, Debezium), Vertex AI or Feature Store
- Compensation: $140,
- Medical, Dental & Vision (employee coverage 100% paid for by Hack The Box)
- 401K with employer match
- Em…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).