Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption

230 · Grab · July 10, 2026, 5:30 a.m.
Summary
This blog post details Grab's transition from using Hive Parquet to Apache Iceberg for managing their extensive data lake. The article discusses architectural challenges faced with increasing data volumes, the benefits of Iceberg adoption, and the creation of a UnifiedSparkCatalog to manage mixed table formats efficiently. It highlights performance improvements, cost reductions, and lessons learned during the migration process. Additionally, Grab is open-sourcing the UnifiedSparkCatalog to contribute to the wider data engineering community.