Iceberg catalogs explained: REST, Glue, S3 Tables, Polaris, Nessie

By Arshad Ansari

"Iceberg catalog" sounds like a big piece of infrastructure: a metadata service, a governance layer, something with a vendor attached. The core job is much smaller. A catalog keeps one pointer per table, to the table's current metadata.json, and changes that pointer atomically when someone commits. Everything else about the table lives in files.

Once you see that pointer, the choice between REST, Glue, S3 Tables, Polaris and Nessie gets easier. They differ in what they add around the pointer, and in which engines can reach them. I have not run Iceberg in production. To show the catalog concretely, I built a table on my laptop with PyIceberg 0.12.0 and DuckDB 1.5.6, using a SQL catalog on SQLite and then on a throwaway Postgres 16. The rows below are from those runs.

What is an Iceberg catalog?

An Iceberg table is a folder of Parquet data files plus a tree of metadata. The top of the tree is a metadata.json holding the schema, the partition spec and the list of snapshots. Every commit writes a new metadata.json. Nothing is edited in place. So after 50 commits there are 51 of them, and something has to say which one is current.

That something is the catalog. Here is the entire state of PyIceberg's SQL catalog after I created one table and appended 50 batches to it. The catalog database held two tables:

iceberg_tables
iceberg_namespace_properties

and iceberg_tables held one row:

catalog_name                lab
table_namespace             shop
table_name                  events
metadata_location           .../events/metadata/00050-c499c071-....metadata.json
previous_metadata_location  .../events/metadata/00049-bb259239-....metadata.json
iceberg_type                TABLE

That is it. One row, two file paths. The table's schema, its 50 snapshots and its file lists are all in the 50 KB metadata.json that row points to. The catalog holds a few hundred bytes.

A commit is a compare-and-swap

When a writer commits, it writes new data and metadata files, then asks the catalog to move the pointer. The SQL catalog does it with one statement. Simplified from PyIceberg's source:

UPDATE iceberg_tables
SET metadata_location          = '.../00051-....metadata.json',
    previous_metadata_location = '.../00050-....metadata.json'
WHERE table_namespace   = 'shop'
  AND table_name        = 'events'
  AND metadata_location = '.../00050-....metadata.json';  -- the version I started from

If another writer committed in between, the WHERE matches no row, and PyIceberg raises CommitFailedException. I tested it on Postgres: two handles on the same table, both appending. The second append conflicted, PyIceberg retried 108 ms later on top of the new state, and both rows landed.

That check is why the catalog exists. Object storage has no cheap way for two writers to agree on which file is "current". The catalog gives them one place to take turns.

After a delete, a schema change, an append, a rewrite and a snapshot expiry, the same row pointed at 00055-....metadata.json, with 00054 as the previous one. Every operation moved the pointer once.

Types of Iceberg catalog

All catalogs do the pointer swap. They differ in protocol, in where they run, and in what they add on top.

CatalogKindRunsAdds
SQL / JDBC catalogDatabase tableYour Postgres, MySQL or SQLiteNothing beyond the pointer
Hive MetastoreMetastoreSelf-hosted serviceFits existing Hive and Spark setups
AWS Glue Data CatalogMetastoreManaged by AWSShared with Athena, EMR and other AWS services
Amazon S3 TablesREST endpoint on table bucketsManaged by AWSAutomatic compaction
Apache PolarisRESTSelf-hosted, or as a managed serviceAccess control, credential vending
LakekeeperRESTSelf-hosted, written in RustAccess control, credential vending
Project NessieREST, Git-styleSelf-hostedBranches, tags, commits across tables
Unity CatalogIts own API; Iceberg REST support varies by editionDatabricks, or OSS self-hostedGovernance, Delta first

Current versions as of 1 October 2026: Polaris 1.8.0, Lakekeeper 0.13.6, Nessie 0.108.8, Unity Catalog OSS 0.6.0, PyIceberg 0.12.0, Iceberg 1.12.0.

The REST catalog is the one to understand. It is an HTTP API defined in the Iceberg project itself. Any server that implements it works with any engine that speaks it. That breaks the old pattern where each engine needed a plugin for each catalog. It also lets the catalog do things a database row cannot: check who is asking, and hand the engine short-lived storage credentials for just that table. Polaris began at Snowflake. Lakekeeper is a smaller, independent implementation. Nessie predates the REST spec and adds Git-style branching on top.

Amazon S3 Tables are worth a separate note. AWS runs the catalog, exposes an Iceberg REST endpoint and compacts the tables for you. As of September 2026 they support Iceberg v3 data types. Supabase's Analytics Buckets, a public alpha, are built on them.

PyIceberg 0.12 ships catalog modules for REST, SQL, Hive, Glue, DynamoDB, BigQuery metastore and an in-memory catalog for tests.

Your engines choose the catalog

The best catalog on paper is useless if one of your engines cannot reach it. Check every engine before you pick.

  • DuckDB. Its May 2026 Iceberg post lists Glue, S3 Tables, Polaris, Lakekeeper and Nessie as catalogs it can attach, and adds writes: MERGE INTO, UPDATE and DELETE. My SQL-catalog table is not on that list. I read it by passing iceberg_scan the metadata path straight from the catalog row.
  • Postgres with pg_lake. Postgres is its own catalog, and it can also create tables in, or attach tables from, a REST catalog (docs).
  • Snowflake. It reads external Iceberg tables through catalog integrations. One for Snowflake Postgres went generally available on 14 July 2026, exposing its tables as read-only Iceberg tables.
  • Spark and Trino have long supported REST, Hive, Glue, JDBC and Nessie catalogs.

REST is the common ground. If you expect more than two engines, start there.

Which Iceberg catalog for a small team

  1. One writer, one engine, one machine or bucket: a SQL catalog on SQLite or on a Postgres you already run. Or no catalog, as below. Fifty commits through a SQLite catalog took 0.84 seconds on my laptop and 0.99 seconds through Postgres. The catalog was never the slow part.
  2. On AWS: S3 Tables if you want compaction handled, Glue if your tables already live in plain S3 buckets. Both are managed, and both are reachable from DuckDB.
  3. Self-hosted, several engines: a REST catalog. Lakekeeper or Polaris, backed by a Postgres you already back up.
  4. You want branches (test a backfill on a branch, then merge it): Nessie.
  5. Databricks or Snowflake at the centre: use the platform's own catalog. Fighting it costs more than it saves.

Whatever you pick, the catalog database is now critical. Lose it and you lose the pointers. The files survive, but you have to work out which metadata.json was current. Back it up like any production database.

Can Postgres be the Iceberg catalog?

Yes, in three ways. Iceberg's JDBC catalog and PyIceberg's SQL catalog use two ordinary tables in any Postgres database, exactly as shown above. Several REST catalogs, including Lakekeeper and Polaris, can use Postgres as their backing store, adding the REST API in front. And pg_lake makes Postgres the catalog for the Iceberg tables it writes, through an iceberg_tables view with the same columns as my SQL catalog, so Spark and PyIceberg can find them.

If Postgres is already your most reliable database, it is a sensible home for the pointers. If you are weighing Postgres for the analytics itself, DuckDB vs Postgres covers that split.

Do you need a catalog at all?

To read, no. DuckDB will read a table from its metadata.json directly:

SELECT count(*) FROM iceberg_scan('warehouse/shop/events/metadata/00055-....metadata.json');
-- 34333

Give it only the table folder and it refuses:

Failed to read iceberg table. No version was provided and no version-hint
could be found, globbing the filesystem to locate the latest version is
disabled by default as this is considered unsafe ...

With SET unsafe_enable_version_guessing = true it picked the newest metadata.json and returned the right count. DuckDB is right to call that unsafe: a reader could pick up a metadata file from a commit that lost its race and was never made current. Some writers also keep a version-hint.text file in the metadata folder, naming the current version. PyIceberg's SQL catalog did not write one in my test.

To write, yes, unless there is exactly one writer and it never overlaps with itself. The catalog's compare-and-swap is the only thing stopping two writers from both believing they made the latest commit.

So a reasonable small setup is a catalog for the writer, and readers that read the metadata.json path you publish, or attach to the same catalog. And if you have one writer, no deletes and no renames, ask whether you need Iceberg at all. Plain Parquet files, which Parquet vs CSV covers, may be enough.

So: which Iceberg catalog?

The catalog is a pointer and a lock. Pick it by two questions. Can every engine you use reach it? And who runs it, backs it up and gets paged when it is down? For most small teams the answer is the managed catalog of the cloud they are already on, or a SQL or REST catalog on the Postgres they already run.

Related: Iceberg vs Parquet asks whether you need a table format at all, and Postgres and Iceberg shows Postgres acting as the catalog.

Local-First Analytics covers the layer underneath: Parquet files, layout and partitioning, and the point at which a table format starts to pay. Chapter 1 is free to read; the rest is on Amazon.

If you are setting up a lakehouse and want a second opinion on the catalog before engines depend on it, a Data Platform Audit is a week and a written roadmap you keep.

Common questions

What is an Iceberg catalog?
An Iceberg catalog is the service that knows, for each table, where its current metadata.json file is, and that changes that pointer atomically when a writer commits. It also lists namespaces and tables. That is nearly all it has to do: the schema, snapshots and file lists live in metadata files in storage, not in the catalog. In a SQL catalog the whole thing is one row per table holding the current and previous metadata location, and a commit is an UPDATE that only succeeds if the pointer has not moved since the writer read it.
What types of Iceberg catalog are there?
Three families. REST catalogs implement the Iceberg REST spec over HTTP: Apache Polaris, Lakekeeper, Nessie and Amazon S3 Tables' endpoint are examples. Database-backed catalogs keep the pointer in a table: the JDBC catalog in Java and the SQL catalog in PyIceberg, on Postgres, MySQL or SQLite. Metastore catalogs reuse an existing service: Hive Metastore and the AWS Glue Data Catalog. PyIceberg 0.12 also ships DynamoDB, BigQuery metastore and in-memory catalogs. REST is the one most engines now speak.
Which Iceberg catalog should a small team use?
Pick the one every engine you use can talk to, then the one you do not have to run. On AWS, S3 Tables or Glue: both are managed, and S3 Tables also compacts for you. Self-hosted with several engines, a REST catalog such as Lakekeeper or Apache Polaris, backed by a Postgres you already run. With one writer and one reader, a SQL catalog on SQLite or Postgres is enough, and you may not need a catalog at all. Choose Nessie if you specifically want Git-style branches across tables.
Can Postgres be an Iceberg catalog?
Yes. Iceberg's JDBC catalog and PyIceberg's SQL catalog both store their state in two ordinary tables, iceberg_tables and iceberg_namespace_properties, in any Postgres database. Each table is one row pointing at its current metadata.json, and commits are a compare-and-swap on that row. Several REST catalogs, such as Lakekeeper and Polaris, can also use Postgres as their backing store. And the pg_lake extension makes Postgres the catalog for the Iceberg tables it creates, readable by Spark and PyIceberg through their JDBC or SQL catalogs.
Do I need a catalog to read Iceberg tables?
Not to read one. DuckDB's iceberg_scan accepts the path to a metadata.json file and reads the table as of that file, with no catalog involved. Given only the table folder, it needs a version-hint file or an explicit opt-in to guess the latest version, which DuckDB disables by default because guessing is unsafe while writers are active. You do need a catalog to write safely, because the catalog's atomic pointer swap is what stops two writers from overwriting each other's commits.

Get new posts by email

Data engineering notes like this one — what breaks and what it costs, in production.

What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.

Want the whole playbook?

If this was useful, the long version is my book. Local-First Analytics — 313 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.

Get the book

Not ready to buy? Read chapter 1 free — the whole chapter, no email required.

Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.