The Open Lakehouse is an open, vendor-neutral data architecture that combines the low-cost storage of a data lake with the transactions, governance, and performance of a data warehouse. It is built on open table formats (Delta Lake, Apache Iceberg, Apache Hudi), open catalogs (Unity Catalog, Apache Polaris), open compute engines (Apache Spark, Trino, DuckDB, Flink), and open ML/AI tooling (MLflow).
The architecture was pioneered by Databricks in the 2020 paper "Lakehouse: A New Generation of Open Platforms" and has since become the default reference design for data and AI platforms across the industry.
Head over to openlakehouse.io to access tutorials, videos, and educational content for all things open-lakehouse.