Sr. Data Platform Engineer
Listed on 2026-08-03
-
Software Development
Data Engineering, Python
About Us
Aarden is a land intelligence platform that helps landowners, investors, and developers figure out what a piece of land can actually be used for, and how to market it. We turn messy parcel, infrastructure, market, community, and ecological data into clear, bankable answers for land-dependent assets. Our goal is to become the default decision layer for land: helping physical projects start in places where they can be built and supported for decades.
We’ve built out a suite of data products to support that goal — pipelines, databases, and AI/ML models that power our maps and In-app agents. We have strong product-market fit, and we re now focused on augmenting our data systems. That’s where you come in.
The roleWe re looking for a Product-Focused data engineer to maintain and evolve our geospatial data pipeline. Our core data asset is a unique blend of property data, geospatial data, and AI-native derived data. Alongside advocating for data excellence, you’ll be empowered to opportunistically contribute to our user-facing product.
What you’ll doPipeline modernization
- Continue our migration of pipeline orchestration to Prefect
- Own day-to-day operations of our data infrastructure
- Extend the pipeline to include new data sources and transformations
- Maintain, expand and optimize our postgres database and Iceberg datalake
- Create the connective tissue for data at Aarden
Cross-team integration
- Partner with product on new feature-driven datasets.
- Collaborate with the ML/analytics team to close the loop: anomaly detection → ticket → fix → validation → promotion to production
- Develop cross-team tooling/infra to keep Git Hub, Notion, Linear, and Slack connected so pipeline issues, docs, and fixes stay linked
Observability & AI-agent readiness
- Implement run-over-run data observability (row counts, key column distributions) to catch anomalies and bugs
- Expose accuracy/quality metrics as first-class artifacts so changes can be evaluated automatically, by a human or an agent
- Write and maintain AI-context documentation (schema docs, pipeline architecture, known patterns/quirks, "what not to do")
Must-have
- Have strong Python skills & are comfortable with PySpark or similar distributed data processing
- Have a strong sense of how the data you working with impacts the end-user
- Are curious and excited about AI and the impact it can have on our ways of working as developers
- Have experience with geospatial data (Geo Parquet, PostGIS, Apache Sedona, or similar)
- Have worked with table formats like Apache Iceberg and lakehouse architectures
- Have worked on workflow orchestration (Prefect, Airflow, Dagster, or similar)
- Are comfortable working in a git-based, CI-friendly workflow
Strongly preferred
- Have worked in full-stack environments, where your work can directly impact the application layer
- Have experience with Apache Sedona or other cloud spatial-compute platforms
- Have built observability/logging layers for data pipelines (not just app services)
- Have experience with property, parcel, real estate, or land data specifically
Nice To have
- Have experience using AI agents to improve data architecture in a real production codebase
- Experience in real estate/land, energy, forestry, or agriculture tech
Languages: Python and SQL. Type Script/Node is a plus for our Application layer and AWS ingest paths.
Orchestration & compute
- Prefect 3 for pipeline orchestration (YAML/config-driven flows, retries, logging)
- Coiled for elastic EC2 workers on GDAL-heavy and batch Python jobs
- Wherobots (managed Apache Sedona / PySpark) for large-scale spatial joins, parcel ingest, and lakehouse work
Data lake & formats: Apache Iceberg, Cloud-Optimized GeoTIFF (COG), and PMTiles. Queried with PySpark and DuckDB.
Geospatial: GDAL, rasterio, Geo Pandas, and tippecanoe. Large-scale spatial work runs on Sedona/Spark via Wherobots.
Databases & serving: PostgreSQL + PostGIS (and pgvector on the app side) as the production store.
Working at AardenAarden is a high-trust, high-output team. We striving to be intentional about our team growth. This allows us to test the outer boundaries of our individual capabilities, while also going deeper on…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).