Senior Software Engineer — Lakehouse Systems
Listed on 2026-09-09
-
Software Development
Data Engineering
Senior Software Engineer — Lakehouse Systems
Location:
Mountain View, CA — On-site
Granica builds AI infrastructure for enterprises operating massive data environments.
Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.
Granica’s products include:
Crunch — continuous optimization for enterprise lakehouse data
Myelin — stateful infrastructure for long-running AI agents
Large Tabular Models — foundation models designed for enterprise tables
Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.
Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.
About the RoleGranica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI.
You will work on the core infrastructure behind Crunch , Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, transaction semantics, table maintenance, object-store-backed storage layouts, file-level optimization, and lakehouse cost/performance across petabyte- and exabyte-scale environments.
You will own core systems that directly affect customer infrastructure cost, query performance, table reliability, and the operational health of large lakehouse environments.
This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, table formats, metadata systems, storage layout, query performance, and AI infrastructure.
You will work on lakehouse systems involving Apache Iceberg, Delta Lake, Apache Hudi, Parquet, ORC, cloud object stores, and query engines such as Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments.
What You’ll DoBuild metadata and transaction systems for large-scale tabular datasets
Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
Implement algorithms in compression, representation, layout optimization, and data efficiency
Contribute to open-source or publish research when appropriate
Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
Hands-on experience with columnar formats such as Parquet or ORC
Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
Strong programming skills in Java, Scala, Go, Rust, C++, or similar…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).