Back to Blog

Iceberg, Delta, Hudi: Pick One in 2026 and Move On

Prateek SinghFebruary 12, 20265 min read84 views
Iceberg, Delta, Hudi: Pick One in 2026 and Move On

The table-format wars are functionally over. Iceberg won on interop. Delta won on installed base. Hudi won on streaming upserts. The decision tree for a new project in 2026 is shorter than the comparison-blog industry wants you to believe.

The war that ended

For five years, choosing a lakehouse table format was a high-stakes, high-confusion decision. Three open formats — Apache Iceberg, Delta Lake, Apache Hudi — each with different metadata layouts, different write semantics, different vendor ecosystems, different decision frameworks pitched by different blog posts on different vendor websites. Picking wrong meant a multi-quarter rip-and-replace.

That war ended in 2024. The endgame move was Databricks' acquisition of Tabular — the company founded by Iceberg's original creators at Netflix — for over $1B (per multiple sources). That acquisition signaled two things at once. First, that Iceberg's momentum was undeniable enough that the company most committed to Delta Lake decided to buy its way into the Iceberg ecosystem rather than keep fighting it. Second, that the future was interoperability rather than format dominance. Databricks immediately built UniForm, which writes Delta tables in a way that exposes them as Iceberg tables to any Iceberg reader. Snowflake added native Iceberg read and write. The format walls came down.

What remained was a much shorter decision tree.

The decision tree

Table-format decision tree — 2026 Three questions. Pick one and move on. Q1: Streaming upserts at >1k records/sec? YES → Apache Hudi copy-on-write or merge-on-read NO Q2: Already on Databricks? YES → Delta Lake + UniForm for Iceberg interop NO → Apache Iceberg the 2026 default Why this is the whole tree: • Iceberg has multi-engine reach (Spark, Flink, Trino, Snowflake, BigQuery, DuckDB) • Delta is operationally cheapest if Databricks Runtime is already paid for • Hudi's record-level upserts beat both for streaming workloads

Iceberg won the interop war

The reason Iceberg is the new default is multi-engine support. By 2026, Iceberg can be read and written natively by Spark, Flink, Trino, Presto, Snowflake, BigQuery, DuckDB, ClickHouse, and a long tail of smaller engines. Delta Lake reaches a similar surface only via UniForm or via Databricks-specific runtime. Hudi has the smallest engine surface of the three.

That matters because greenfield architectures in 2026 are not single-engine. The same data needs to be queried by an interactive engine (Trino), a streaming engine (Flink), a notebook engine (DuckDB), and possibly a warehouse (Snowflake) — all without copying. Iceberg is the format every engine has agreed to read. Iceberg's vendor-neutral governance under the Apache Software Foundation, combined with the open architecture of catalog implementations like Polaris and Nessie, means the format is not going to get rugged out from under anyone.

Delta still has the installed base

The case for Delta in 2026 is operational: it is used by over 60% of Fortune 500 companies and 10,000+ organizations, mostly through Databricks' customer base. If you are already on Databricks, switching off Delta is paying real money to solve a problem you don't have. UniForm gives you Iceberg-compatible reads from your existing Delta tables, which closes the interop gap for most use cases.

The case against Delta for greenfield is that the optimization story still favors Databricks Runtime. Delta on open Spark works, but the most-tuned execution path lives behind a Databricks subscription. For organizations that aren't paying that bill, Iceberg's multi-engine performance is more even.

Hudi for the streaming sliver

Hudi is the right answer for a narrow but real use case: high-frequency streaming upserts. The classic shape is CDC ingestion from a high-write OLTP database into a lakehouse, where you need record-level upsert semantics at thousands of records per second with minimal latency. Hudi's merge-on-read tables and built-in indexing infrastructure handle this pattern more cleanly than either Iceberg or Delta.

Outside that sliver, Hudi is the third place. The community is smaller, the engine support is narrower, and the operational tooling lags. If you don't need the streaming-upsert features specifically, picking Hudi in 2026 is choosing a smaller ecosystem for no compensating benefit.

The interop layer is the real story

The most important quiet shift is that Iceberg has become the interoperability lingua franca, and the other two formats are increasingly defined by their relationship to it. Delta Lake's UniForm exposes Delta as Iceberg. Hudi has shipped native Iceberg-compatibility tooling. Even Paimon and DuckLake — the newer entrants — speak Iceberg. The format wars are over not because anyone won, but because Iceberg became the common protocol everyone agreed to read.

For most teams in 2026, the right move is to stop running comparison spreadsheets, pick the answer your decision tree gives you above, and spend the saved cycles on the things that actually differentiate your platform — query patterns, data modeling, governance, observability, cost. The format itself isn't your moat. The pipeline that lands clean data into it is.

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Data Engineering10 min read

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.

59 views
Read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Data Engineering19 min read

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.

58 views
Read
DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower
Data Engineering15 min read

DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower

DuckDB has no GPU backend of its own, and the GPU engines built for it need an NVIDIA card. gpudb 0.7 is my Apache-2.0 DuckDB extension for the GPU already inside your Mac, and for CUDA too. You write plain DuckDB SQL; the GPU takes a statement only where it has been measured faster than DuckDB on your own machine. On an Apple M4 Max, 19 of 22 TPC-H SF10 queries run on the Metal GPU at 1.06x to 48x with zero rows differing. The one row below parity is printed, not dropped.

78 views
Read