unitycatalog-playground is a complete, interactive environment for running open source Unity Catalog locally on macOS or Linux. It connects Unity Catalog to Apache Spark and Delta Lake, stores catalog metadata in PostgreSQL, stores managed-table data in S3-compatible object storage, and provides ready-to-run marimo notebooks.
Use it to create, write, and query catalog-managed Delta tables—not just start a catalog server.
- Unity Catalog OSS for catalogs, schemas, tables, volumes, and credentials
- Apache Spark + Delta Lake as the local compute and table engine
- PostgreSQL for persistent Unity Catalog metadata
- RustFS for local S3-compatible managed-table storage and credential vending
- marimo notebooks with runnable Unity Catalog and Delta Lake examples
- Local and remote modes for using the bundled catalog or another UC server
This project complements the official Unity Catalog quickstart. The upstream quickstart is the best place to learn the server and APIs. This playground adds the compute engine, durable metadata database, object storage, and notebooks needed for an end-to-end local lakehouse.
The container images used by the playground support both Apple Silicon
(arm64) and Intel (amd64) Macs.
- Docker Desktop with at least 4 GB of memory available to Spark
- Git
just, the project command runner
Install and start the prerequisites with Homebrew:
brew install --cask docker
brew install just
open -a DockerAfter Docker Desktop is running:
git clone https://github.com/open-lakehouse/unitycatalog-playground.git
cd unitycatalog-playground
just uc=local startThe first start builds the notebook image and may take several minutes. When the environment is ready, the command prints a URL like:
http://localhost:2718/?access_token=...
Open that URL, select unitycatalog-delta.py, and run the notebook cells in
order. You will create a Unity Catalog schema and a catalog-managed Delta table,
write sample data through Spark, and query it again.
The local services are available at:
- marimo notebook environment: http://localhost:2718
- Unity Catalog API: http://localhost:8080
- RustFS object browser: http://localhost:9001
- RustFS S3 API: http://localhost:9000
To stop the stack while preserving its metadata and table data:
just uc=local downOne docker-compose.yaml supports two modes:
- Local:
just uc=local ...starts the complete stack: Unity Catalog, PostgreSQL, RustFS, and marimo with Spark and Delta Lake. - Remote:
just ...starts only the notebook environment. ConfigureUC_SERVER_URL,UC_SERVER_PORT, andUC_TOKENin.envto connect it to an external Unity Catalog server.
Use the same mode when starting and stopping the environment:
just uc=local start # build and start the complete local stack
just uc=local logs # follow logs
just uc=local ps # show service status
just uc=local down # stop containers but preserve data
just start # connect the notebooks to a remote UC server
just downRun just init to create .env from .env.example. The
defaults work for the local quickstart. Edit .env to change Spark resources,
ports, PostgreSQL or RustFS settings, the remote UC endpoint, or proxy settings.
For environments that require package proxies:
PYPI_PROXY_URL=https://pypi-proxy.yourcompany.com/simple
MAVEN_PROXY_URL=https://maven-proxy.yourcompany.com/maven2Docker Compose loads .env automatically. Leaving those values blank uses
public PyPI and Maven Central.
marimo-spark joins an external docker network, uc-shared, so it can also
reach a UC server running in a separate compose project. Because the network is
declared external: true, it must exist before Compose starts — otherwise you'll
see network uc-shared declared as external, but could not be found.
just up / just up-detached / just jars run just net first, which creates
the network idempotently. You can also run it on its own:
just net # docker network create uc-shared (no-op if it already exists)By default each notebook resolves its jars (delta-spark, unitycatalog-spark,
hadoop-aws, and their transitive dependencies) from Maven the first time a
Spark session is created. That resolution can take several minutes and needs
network access — and it repeats on a fresh container or whenever the Ivy cache
is cleared.
just jars does that resolution once, up front, and drops the jars into
./spark/jars (bind-mounted to /spark/jars in the container). A notebook that
reads from that directory then starts Spark straight off the local classpath —
no per-session Maven round-trip.
just build # the image must exist first
just jars # resolve + download into ./spark/jars
just up # start the environmentNote:
just jarsresolves throughMAVEN_PROXY_URL(from your.env) when it is set. It uses a proxy-only Ivy resolver (spark/ivysettings.xml), avoiding attempts to reach the firewalledrepo1.maven.organd Spark Packages endpoints first. WhenMAVEN_PROXY_URLis blank, it resolves from public Maven Central.
The downloaded jars are git-ignored (the directory is kept via
spark/jars/.gitkeep), so they never get committed. Re-running just jars is
safe—existing jars are not re-downloaded. The coordinates it fetches live in
the jars_packages variable at the top of the Justfile; keep them
in sync with the spark.jars.packages used in the notebooks.
- Use it for fast, repeatable notebook startup across container rebuilds.
- Use it behind a slow or restricted corporate network.
- Skip it for a quick one-off run; the notebooks resolve dependencies from
Maven automatically when
./spark/jarsis empty.
unitycatalog-delta.pycreates a schema and a catalog-managed Delta table, generates sample data, writes it through Spark, and queries it.metric-views.pydemonstrates Unity Catalog metric views with reusable dimensions and measures.delta-new-in-4.4.0.pyexplores Delta Lake 4.4 features on Apache Spark 4.2.external-access-unitycatalog-delta-managed-read.pydemonstrates reading a remote Databricks Unity Catalog table through external data access and credential vending.
Open the marimo URL printed by just ... start, select a notebook, and run its
cells in order. Markdown cells provide context and do not need to be executed.
Spark runs inside the marimo container and uses the Unity Catalog Spark connector. Unity Catalog stores object definitions in PostgreSQL and vends short-lived credentials for managed-table data in RustFS. This gives local development the same basic separation of catalog metadata, compute, and object storage used by a cloud lakehouse.
just uc=local … starts Postgres 16.3 alongside the UC server (same defaults as
upstream postgres-example.yml).
Hibernate is pointed at it via etc/conf/hibernate.properties.
Catalog metadata lives in the Docker volume uc_postgres_data and survives
just down / restarts. Wipe it with just uc=local clean or
just uc=local down-volumes.
Managed table data lives in RustFS, an S3-compatible
object store, rather than a local file path — so the local stack behaves like a
cloud UC on real object storage (storage-root.tables=s3://uc-warehouse). A
one-shot rustfs-init service creates the bucket, mints a short-lived STS
credential from RustFS, and renders it into the UC server config so UC can vend
it to Spark. (This UC build has no custom-S3-endpoint support yet, so the STS
token is minted from RustFS rather than assumed by UC directly — see
etc/rustfs/bootstrap.sh.)
just uc=local start
just uc=local rustfs-urlOpen http://localhost:9001 and log in with RUSTFS_ACCESS_KEY and
RUSTFS_SECRET_KEY (both default to rustfsadmin) to browse the objects your
notebooks write.
The vended STS credential expires (default 12h). If managed-table reads or writes start failing with credential errors, refresh it — the bucket and its data are left untouched:
just uc=local rotate-creds # re-mint the STS credential + restart UCObject data lives in the rustfs_data volume (wiped by
just uc=local clean or just uc=local down-volumes). Tune the bucket,
credentials, region, and STS lifetime with RUSTFS_*, UC_STORAGE_BUCKET,
S3_REGION, and STS_DURATION_SECONDS in .env.
Stop containers and networks while keeping PostgreSQL metadata and RustFS table data:
just uc=local down
# or: docker compose --profile local-uc downTo reset everything, delete the named volumes and the locally built marimo image:
just uc=local cleanCaution
clean permanently deletes the local Unity Catalog metadata and all managed
table data stored by the playground.
Start Docker Desktop and wait for its engine to become ready:
open -a Docker
docker infoThe just commands create this network automatically. If Compose was invoked
directly or startup was interrupted, recreate it with:
just netSpark downloads Delta Lake, the Unity Catalog connector, Hadoop AWS, and their transitive dependencies the first time a notebook initializes. Pre-download them for subsequent runs:
just build
just jarsGive Docker Desktop at least 4 GB of memory for Spark. You can also lower or
raise SPARK_DRIVER_MEMORY and SPARK_DRIVER_CORES in .env, then restart the
marimo container and notebook kernel.
The local RustFS STS credential expires after 12 hours by default. Refresh it without deleting the bucket or its data:
just uc=local rotate-credsThe helper module at
delta/python/catalog-managed.py provides
reusable utilities for creating and populating catalog-managed Delta tables
with PySpark. The
unitycatalog-delta.py
notebook uses the same helpers interactively.
A simple dataclass representing a rescue animal record.
@dataclass
class Pets:
uuid: str # unique identifier
name: str # pet name
age: int # age in years
adopted: boolGenerates total random Pets records and yields them in batches of batch_size.
pets = generate_pets(batch_size=10, total=100)
first_batch = next(pets) # list[Pets] with 10 records
for batch in pets: # iterate remaining batches
print(batch)Converts a list[Pets] batch into a Spark DataFrame with an explicit schema.
df = pets_to_dataframe(first_batch, spark)
df.show()Schema:
| column | type | nullable |
|---|---|---|
uuid |
string | no |
name |
string | no |
age |
integer | no |
adopted |
boolean | no |
Builds a CREATE TABLE IF NOT EXISTS ... USING DELTA DDL string from a
StructType schema and an optional dict of TBLPROPERTIES.
props = {"delta.feature.catalogManaged": "supported"}
ddl = create_table_ddl("sanctuary.pets", df.schema, props)
print(ddl)Executes the DDL produced by create_table_ddl via spark.sql. Returns an
empty DataFrame on success.
create_table_using_sql("sanctuary.pets", df.schema, props, spark)The following mirrors the flow in the
unitycatalog-delta.py
notebook:
# 1. Create the schema
spark.sql("CREATE SCHEMA IF NOT EXISTS unity.sanctuary")
# 2. Generate pets
pets = generate_pets(batch_size=10, total=100)
first_batch = next(pets)
# 3. Derive the schema from the first batch
df = pets_to_dataframe(first_batch, spark)
props = {"delta.feature.catalogManaged": "supported"}
# 4. Create the table
create_table_using_sql("sanctuary.pets", df.schema, props, spark)
# 5. Write the first batch
df.write.format("delta").mode("append").saveAsTable("sanctuary.pets")
# 6. Write remaining batches
for batch in pets:
pets_to_dataframe(batch, spark).write.format("delta").mode("append").saveAsTable("sanctuary.pets")
# 7. Verify
spark.sql("SELECT COUNT(*) AS total FROM sanctuary.pets").show()Yes. The stack's Unity Catalog, Spark, PostgreSQL, RustFS, and AWS CLI container
images publish linux/arm64 variants. Docker Desktop selects them automatically
on Apple Silicon. The same stack also supports Intel (amd64) Macs.
No. Local mode uses the open source Unity Catalog server and runs entirely in Docker. A Databricks account is needed only for examples that explicitly connect to a Databricks-managed Unity Catalog.
The official repository provides the authoritative Unity Catalog server, clients, APIs, and UI. This community playground builds on that server and adds Spark, Delta Lake, PostgreSQL, S3-compatible storage, credential vending, and interactive notebooks so you can exercise complete table workflows locally.
Yes. Run in remote mode and set UC_SERVER_URL, UC_SERVER_PORT, and
UC_TOKEN in .env:
just startNo. The playground is designed for local development, demonstrations, and experimentation. Use production-grade identity, networking, storage, observability, backups, and secret management for a deployed environment.
Issues and pull requests are welcome. When reporting a startup problem, include your Mac architecture, Docker Desktop version, the command you ran, and the relevant output from:
just uc=local ps
just uc=local logsLicensed under the Apache License 2.0.