Iceberg vs Parquet: do you need a table format?

By Arshad Ansari

"Apache Iceberg vs Parquet" sounds like a choice between two file formats. It isn't. An Iceberg table stores its data in Parquet files. Iceberg is a table format: a set of metadata files that say which Parquet files make up the table right now, what the schema is, and what the table looked like at every past commit.

So the useful question is: do your Parquet files need a table format on top? For many small teams the honest answer is no. A note on where this comes from: I run plain Parquet with DuckDB in production. I have not run Iceberg in production. To get numbers instead of opinions, I built Iceberg tables on my laptop with PyIceberg 0.12.0, PyArrow 25.0.1 and DuckDB 1.5.6, on the local filesystem with a SQLite catalog. Every figure below comes from those runs.

Apache Iceberg vs Parquet: a file and a table

A Parquet file is self-contained. It knows its own columns, types and row groups, and nothing about any other file. A "table" made of plain Parquet is a folder and a naming convention. The engine lists the folder and reads what it finds.

An Iceberg table adds three layers above the same kind of files. After 50 small appends, my test table looked like this on disk:

warehouse/shop/events/
  data/
    00000-0-061edf85-....parquet          50 Parquet files, 450 KB
  metadata/
    00050-c499c071-....metadata.json      the current table state (50 KB)
    snap-616794116452149308-....avro      manifest list: one per snapshot
    ....avro                              manifests: lists of data files + stats

Reading goes top down. The catalog points at the current metadata.json. That file holds the schema, the partition spec and the list of snapshots. Each snapshot points at a manifest list, which points at manifests, which list the data files with per-column statistics. A commit writes new files and then swaps the catalog pointer in one step. That swap is what makes a commit atomic.

The format is versioned. The Iceberg spec has versions 1 to 3 adopted and version 4 in development. Version 3 adds a variant type, nanosecond timestamps, geospatial types, default values, deletion vectors and row lineage. The Java reference implementation reached 1.12.0 on 30 September 2026.

What plain Parquet already gives you

My book's position is that plain Parquet plus DuckDB is enough for a lot of real work, and I still hold it. Each file carries its own schema, so DuckDB can read a folder whose files gained a column over time:

SELECT * FROM read_parquet('data/*.parquet', union_by_name = true);
-- files written before the new column return NULL for it

That handles the most common change, adding a column. The book is just as plain about where it stops. Renames, type changes and dropped columns still break consumers. There is no versioning, no time travel and no atomic commit across many files. Its rule: for single-writer, append-only work, plain Parquet. For multi-writer warehouses, reach for a table format.

What Iceberg adds, on a 50-commit table

I appended 50 batches of 1,000 rows, then did the things plain Parquet handles badly. Each time I compared DuckDB's iceberg_scan on the table with a plain read_parquet glob over the same data folder.

Deletes that stay deleted. I deleted every row where kind = 'view'. PyIceberg rewrote the affected files. The data folder went from 50 files to 100, because the old files stay until someone cleans them up.

iceberg_scan:          33,333 rows   (correct)
read_parquet glob:     83,333 rows   (old and new files, double-counted)

With plain Parquet, a delete means rewriting files and making sure no reader ever sees both versions. Iceberg does this by listing only the live files in the current snapshot.

Renames that don't lose data. I renamed amount to amount_inr, added a channel column and appended 1,000 more rows. Iceberg tracks columns by numeric field ID, not by name, so old files still map to the renamed column.

iceberg_scan:                        34,333 rows, 34,333 with amount_inr
read_parquet glob, union_by_name:    84,333 rows,  1,000 with amount_inr

The plain glob saw amount_inr only in the one new file. Every older value had silently become NULL.

Time travel. Reading the table as of its first snapshot returned exactly 1,000 rows, the first batch. Every commit stays readable until you expire it.

Several writers. I opened the table twice, as two writers, and appended from both. The second commit found the pointer had moved. The catalog rejected it, PyIceberg retried 108 ms later on top of the new state, and both rows landed. Plain Parquet has no such check. Two writers just write files and hope.

Partition changes without a rewrite. Iceberg also supports hidden partitioning (queries filter on a timestamp, and Iceberg maps that to the day partition itself) and changing the partition scheme for new data while old data keeps its layout. I did not test these. They come from the spec.

What Iceberg costs: files, commits and upkeep

Every commit writes a data file and new metadata. With small, frequent commits, metadata stops being small. After the 50 appends:

FilesSize
Data (Parquet)50450 KB
Metadata (json + manifest lists + manifests)1511,643 KB

The metadata was 3.6 times the size of the data. Then I pushed it: 500 commits of 200 rows each.

commit time, median of first 10:     9 ms
commit time, median of last 10:    202 ms
data files:      500,      856 KB
metadata files: 1,501, ~106 MB

About 100 MB of that was old metadata.json files. Each commit writes a fresh one, and each carries the full snapshot history, so they grow with every commit. Iceberg has table properties for this. With write.metadata.delete-after-commit.enabled = true and write.metadata.previous-versions-max = 10, the same 500 commits left about 12 MB of metadata. Commit time still rose, to 257 ms, because the current metadata.json still listed 500 snapshots.

Small files cost reads too. DuckDB took a median 47.5 ms to sum the 500-file table and 3.2 ms after I rewrote it into one file. On a laptop those are both instant. On object storage, each file is at least one request.

So compaction matters, and it is a job someone runs. In PyIceberg 0.12 I found snapshot expiry but no compaction call, so I compacted by overwriting the table with its own contents, which took 1.8 seconds. Snapshot expiry then removed every snapshot but the current one from the metadata, but every old file was still on disk: 501 data files, about 109 MB. Deleting unreferenced files is a separate step again.

Some platforms run this for you. pg_lake's Iceberg docs say it vacuums Iceberg tables every 10 minutes, compacting files and removing expired ones. Amazon S3 Tables compacts automatically. DuckDB can now write Iceberg, with MERGE INTO, updates and deletes, per its May 2026 post, but that post announced no compaction.

The last cost is the catalog: one more service to run, secure and back up. That deserves its own post.

Iceberg vs Parquet performance

Iceberg does not make Parquet faster. It reads the same files, after reading some metadata first. On my 50-file table, DuckDB took 10 ms through iceberg_scan and 2 ms with a direct glob. On a table this small, the metadata is the overhead.

At scale the balance can flip. Manifests store per-file column statistics, so an engine can skip files that cannot match a filter without opening them. Iceberg also avoids listing directories, which is slow on object storage with many thousands of files. I did not test at that scale, so I can't give you a crossover point. Measure on your own data. The safe summary: Iceberg is about correctness and coordination. Any speed gain comes from pruning and layout, which you also have to maintain.

If you want the speed story for the files themselves, Parquet vs CSV has the numbers.

Iceberg vs Parquet vs Delta Lake

Delta Lake solves the same problem as Iceberg in a different way. Both put Parquet data files under a metadata layer that gives atomic commits, snapshots, time travel and deletes. Delta records each commit as a file in an ordered _delta_log folder beside the data. Iceberg keeps a metadata tree and leans on a catalog for the current pointer. Apache Hudi is a third option, with a focus on upserts and incremental pulls.

The choice is mostly made by your engines, not by the formats:

  • Databricks at the centre: Delta. Even Databricks Lakebase, its managed Postgres, syncs to Delta in Unity Catalog, not Iceberg.
  • Several engines, or AWS or Snowflake at the centre: Iceberg has the wider reach. S3 Tables, Snowflake, Spark, Trino, DuckDB and Postgres extensions such as pg_lake all read it.
  • One engine and one writer: maybe neither.

The current releases, for reference: Iceberg 1.12.0, Delta Lake 4.4.0, Hudi 1.2.1.

When plain Parquet files are enough

Plain Parquet is enough if you can say yes to all five:

  1. One process writes each dataset. Nothing else writes to the same folder.
  2. Data is appended, or whole partitions are rebuilt. You never change individual rows in place.
  3. Schema changes are new columns. No renames, no type changes.
  4. Readers never catch a half-written partition: you write to a temporary path and rename, or readers run after the writer.
  5. Nobody needs to ask what the table held last Tuesday, or can answer it from dated partitions.

My own pipeline's raw layer meets all five, so it stays plain Parquet. Add a table format when one of these turns into a no. The most common triggers I see are row-level deletes for erasure requests, a second engine or team writing to the same table, and a column rename that cannot be avoided.

So: do you need Iceberg?

If your pipeline is one writer appending Parquet that DuckDB or Polars reads, no. You would take on a catalog, a growing pile of metadata and a compaction job, and get guarantees your workload never asks for.

If several writers or engines share a table, if rows must be deleted for real, or if the warehouse needs to read your lake, yes. Plan the compaction and cleanup jobs on day one, not after the first slow query.

This post has two siblings: Postgres and Iceberg covers pg_lake and pg_duckdb, and Iceberg catalogs explained covers the catalog every Iceberg table needs.

Local-First Analytics covers the plain-Parquet side in depth: file layout, partitioning, schema changes and where a table format starts to pay. Chapter 1 is free to read; the rest is on Amazon. Related reading: Parquet vs CSV: what you actually save, local-first analytics in practice with DuckDB and Parquet, and is DuckDB safe for production?

If you are weighing a lakehouse and want a second opinion on whether your workload needs one, a Data Platform Audit is a week and a written roadmap you keep. If plain Parquet is enough, that is what the roadmap will say.

Common questions

Is Apache Iceberg a file format?
No. Apache Iceberg is a table format: a specification for metadata files that say which data files make up a table, what its schema is, how it is partitioned and what it looked like at each past commit. The data itself usually sits in Parquet files (ORC and Avro are also allowed). So an Iceberg table is a folder of ordinary Parquet files plus a tree of metadata files (a metadata.json, manifest lists and manifests) and a catalog entry that points at the current metadata.json. Asking "Iceberg or Parquet" is like asking "Postgres or its data pages": one is built on the other.
Is Iceberg faster than Parquet?
Not by itself. An Iceberg table reads the same Parquet files, plus a few metadata files first. On a small table that makes it slightly slower: in a laptop test, DuckDB 1.5.6 read a 50-file table in 10 ms through Iceberg and 2 ms by globbing the Parquet files directly. Iceberg wins at scale for two reasons: it can skip files using column statistics stored in its manifests, and it avoids listing directories on object storage. Its real advantage is correctness, not speed: after a delete, the raw glob in the same test returned 83,333 rows where the table held 33,333.
What is the difference between Iceberg, Parquet and Delta Lake?
Parquet is a columnar file format. Iceberg and Delta Lake are both table formats that sit on top of Parquet files and add atomic commits, snapshots, time travel, schema evolution and row-level deletes. They differ in how they record commits: Iceberg keeps a tree of metadata files and relies on a catalog to point at the current version, while Delta Lake keeps an ordered log of commit files in a _delta_log folder next to the data. In practice the choice follows your engines: Databricks shops lean to Delta, while Iceberg has the wider spread across Snowflake, AWS, DuckDB, Spark, Trino and Postgres extensions such as pg_lake.
When are plain Parquet files enough?
Plain Parquet is enough when one process writes, data is appended or whole partitions are rebuilt rather than individual rows changed, schema changes are limited to adding columns, and readers can tolerate seeing a partition mid-write or you write to a temp path and rename. That covers many small teams' analytics: daily batch jobs writing Parquet that DuckDB or Polars reads. Reach for a table format when you need row-level deletes (for example GDPR erasure), column renames, several writers or engines on one table, or an audit of what a table held on a past date.
What is the small files problem in Iceberg, and how does compaction fix it?
Every Iceberg commit writes at least one new data file plus new metadata files, so frequent small writes leave a table made of many tiny files and a growing pile of metadata. In a laptop test, 500 commits of 200 rows produced 500 data files totalling 856 KB and about 106 MB of metadata, and commit time rose from 9 ms to 202 ms. Compaction rewrites many small data files into a few large ones; snapshot expiry and orphan-file cleanup then remove the old files. Someone has to run these jobs, unless your platform does it, as pg_lake's autovacuum and Amazon S3 Tables do.

Get new posts by email

Data engineering notes like this one — what breaks and what it costs, in production.

What breaks and what it costs — pipelines, warehouse bills, and the failures that only show up in production. A few a month, never padded to hit a schedule. No sequence, no pitch deck. Reply 'stop' once and you're off — it reaches me, not a queue.

Want the whole playbook?

If this was useful, the long version is my book. Local-First Analytics — 313 pages, runnable code for every chapter — is the full build: DuckDB, Parquet and Arrow, from install to production. On Amazon, and chapter 1 is free to read here.

Get the book

Not ready to buy? Read chapter 1 free — the whole chapter, no email required.

Rather talk it through? Book a free 30-minute call. No slot that suits your time zone? Email info@hikmahtech.in.