L2 – Senior Iceberg DBA/Lakehouse Operations Engineer Remote
Plano, Collin County, Texas, 75023, USA
Listed on 2026-07-19
-
IT/Tech
Data Engineering, Data Warehousing
L2 – Senior Iceberg DBA / Lakehouse Operations Engineer
Location:
Remote work accepted from anywhere in US
Duration:
Long Term
Job Description:
• 4–6 years of experience in Big Data / Data Operations / DBA roles
• Minimum 1+ year of experience with Apache Iceberg or similar table formats (Hive/Delta/Hudi)
• 4+ years of experience with Cloudera ecosystem (CDP)
• Hands-on experience with:
Iceberg table operations and maintenance, Spark SQL, Hive, or Impala
• Experience in:
Production support and incident handling, Monitoring, troubleshooting, and operational support
• Apply established data modeling and Lakehouse standards in day-to-day operations
• Support:
Table structuring, Partition alignment with ingestion patterns
• Assist in maintaining consistency of datasets across Bronze/Silver/Gold layers
Required Skills:
• Strong hands-on experience with Apache Iceberg and/or Hive-based data lakes
• Understanding of data modeling concepts (normal forms) and modern Lakehouse patterns (Medallion architecture)
• Expertise in:
Table-level optimization and performance tuning, Large-scale data management (TB/PB scale)
• Experience with:
Spark SQL, Hive, Impala, NiFI, Trino
• Strong understanding of:
Partitioning strategies, File formats (Parquet/ORC), Distributed query processing
Preferred
Skills:
• Experience with:
Hive-to-Iceberg or Teradata-to-Iceberg migration, Cloudera CDP (CDE/CDW)
• Familiarity with:
Cloud platforms (AWS, Azure)
• Scripting/automation (Python, Shell)
What You’ll Work On:
• Enterprise-scale Iceberg Lakehouse platform supporting multiple applications
• Large-scale data modernization initiatives
• Performance optimization and stability of mission-critical analytical workloads
Why This Role Matters:
• Ensures data correctness and performance for downstream analytics and business-critical reporting
• Enables successful modernization from legacy platforms to Iceberg
• Maintains high availability and reliability of the enterprise data layer
Job Summary:
We are seeking a highly skilled Iceberg DBA / Lakehouse Operations Engineer to own the reliability, performance, and operational integrity of the Iceberg data layer powering enterprise analytics and business-critical applications. This role operates in a large-scale, multi-engine Lakehouse environment, supporting workloads across Spark, Hive, and Impala, and plays a key role in enterprise data modernization initiatives (Hive and Teradata → Iceberg).
The ideal candidate brings deep expertise in Iceberg table operations, metadata management, and query performance optimization, ensuring consistent, high-performance data access across platforms in a cloud-based environment. This role is critical to ensuring data accuracy and performance—any degradation directly impacts downstream reporting, analytics, and business-critical decision-making.
Key Responsibilities:
Iceberg Data Layer Ownership & Operations:
• Own day-to-day operations of Apache Iceberg tables supporting multiple enterprise applications
• Ensure data reliability, consistency, and availability across all Lakehouse workloads
• Maintain operational integrity for datasets at multi-terabyte to petabyte scale
Advanced Table Management & Optimization:
• Execute advanced Iceberg table maintenance and optimization strategies:
Compaction (minor/major) and small file mitigation, Snapshot expiration and metadata compaction to control metadata growth, Orphan file cleanup (vacuum) to maintain storage efficiency
• Optimize data layout and performance through:
File size tuning and distribution strategies, Partition evolution and pruning optimization, Clustering and ordering techniques (e.g., Z-ordering or similar patterns)
Data Modeling Standards & Lakehouse Design Alignment:
• Support and enforce data modeling best practices aligned with:
Normalized data structures (3NF) for source-aligned datasets, Medallion architecture (Bronze / Silver / Gold layers) for curated data flows
• Ensure Iceberg table design aligns with:
Data ingestion patterns (raw vs curated layers), Downstream consumption and performance requirements
• Assist in structuring datasets to balance:
Data integrity and normalization, Query performance and analytical…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).