AWS Lakehouse Data Engineer
Listed on 2026-09-04
-
Software Development
Data Engineering, AWS
Job Family:
Software Development & Support
Travel Required:
None
Clearance Required:
Ability to Obtain Public Trust
We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.
This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.
Build and Operate Data Pipelines (Batch and Streaming)
Build and Operate Data Pipelines (Batch and Streaming) Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners. Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning. Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
Design and implement a Delta Lakehouse Data Platform Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization. Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet. Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate. Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage. Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.
Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services. Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability. Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling. Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.
AWS Automation, CI/CD, and Operations Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines. Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies. Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures. Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
Cross-Team Collaboration and DocumentationCross-Team Collaboration and Documentation Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines. Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines. Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.
What You Will NeedBachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).