Back to Blog

DuckDB Ate the Modern Data Stack

Prateek SinghFebruary 17, 20266 min read94 views
DuckDB Ate the Modern Data Stack

An embedded analytical engine with no servers, no cluster, no migration cost just quietly displaced Spark for small data and Snowflake XS for medium data. MotherDuck closed Series B at a $400M post-money. Here's the part everyone undercounts.

The thing nobody priced in

For a decade, the assumption underneath every "modern data stack" deck was that analytical compute had to live behind a network boundary. You shipped your data to Snowflake, BigQuery, Redshift, or a Spark cluster, paid for warm warehouses or autoscaled clusters, and accepted the latency tax because the alternative was a single-machine Postgres that fell over at 100 GB.

That assumption was correct in 2014. It stopped being correct around 2022. The reason is the specific shape of what DuckDB did: it took a single-process columnar engine — no servers, no cluster, no orchestration — and made it fast enough that the crossover point where you actually needed a warehouse moved up by roughly two orders of magnitude. Suddenly the workload that justified a Snowflake XS warehouse and an Airflow DAG was a one-line duckdb -c "SELECT ... FROM read_parquet('s3://...')".

The market noticed. MotherDuck, the commercial DuckDB cloud, raised a $52.5M Series B led by Felicis at a $400M post-money valuation — $100M total funding, with a16z, Madrona, Amplify Partners, Altimeter, Redpoint, and Zero Prime all in (PR Newswire). DuckDB itself is open-source and free, so the funding signal is about something downstream: investors believe the workload split between "embedded DuckDB on a laptop or container" and "DuckDB-as-a-managed-service" is going to be enormous, and someone has to host the second half.

The crossover chart

The clearest way to see what DuckDB ate is to plot the workload-size axis against the right tool to use. Pre-DuckDB, the chart had two regions: Postgres for <10 GB, warehouse for everything above. Post-DuckDB, there are three regions, and the middle one is enormous.

Workload-size to engine map — 2026 Where the right-tool boundary moved when DuckDB matured. Pre-DuckDB (2019) Postgres < 10 GB Warehouse cluster (Snowflake / BigQuery / Spark) 10 GB to multi-PB Post-DuckDB (2026) Postgres OLTP only DuckDB (embedded or MotherDuck) 1 GB to 500 GB — the new middle Warehouse cluster > 500 GB The 1 GB-500 GB band is most of analytical work in the world. DuckDB now owns it.

That middle band — 1 GB to 500 GB — is most of analytical work in the world. Customer dashboards, finance close, ad-hoc data-science notebooks, ETL transforms, ML feature builds, internal reporting. None of it actually needed a cluster. The cluster existed because Postgres was the only single-machine alternative and Postgres was bad at columnar scans. Take Postgres out of the picture and the cluster's whole reason for existing in the small-to-medium range disappears.

What DuckDB actually killed

Three workloads died loudly between 2023 and 2026:

  1. Spark for <500 GB. If your dataset fits in a modern laptop's RAM (and most do), Spark's per-job startup cost, JVM overhead, and operational tax aren't worth paying. DuckDB on a single c6i.4xlarge crushes a small Spark cluster on the same workload at a fraction of the operational complexity. Databricks knows this — it's why they've been quietly building Photon as a single-node fast path and why Spark Connect is their answer to the embeddability question.
  2. Snowflake XS for ad-hoc. An XS warehouse costs $2/credit and burns a credit per hour even when idle (with the auto-suspend grace). For analysts who want to interrogate a 50 GB Parquet file three times an afternoon, that's pure waste. They're now opening DuckDB in a Jupyter notebook and pointing it at the same S3 path.
  3. The "transformation layer" of the modern data stack. dbt-on-Snowflake is being quietly displaced by dbt-on-DuckDB for any project where the source data is Parquet on object storage. The transforms run locally in CI, the artifacts land in a warehouse only if you actually need shared concurrent SQL access on top.

The single best demonstration of how far the engine has come is the memory profile. DuckDB now keeps peak memory under 2.5 GB even on 2 TB datasets, and partitioning a 140 GB dataset into smaller files cuts peak memory 8× to 160 MB (per the DuckDB Ecosystem Newsletter, February 2026). That is not a "small-data" engine. That is a small-machine engine running large-data workloads, which is a different and far more interesting category.

Where DuckDB doesn't go

Honest accounting: DuckDB is not the answer for everything, and three categories of workload are still warehouse-shaped:

  • High-concurrency BI dashboards. DuckDB is single-process. If you need 200 concurrent analysts pinging the same tables, you still need a multi-tenant warehouse (or you put MotherDuck or DuckLake or some shared layer in front of DuckDB).
  • Multi-PB scans. Above ~5 TB the constants flip and a distributed engine wins.
  • Strong cross-team governance. A warehouse with a unified RBAC model and audit log is still the simplest answer if compliance matters.

Everything else? The right answer in 2026 is "try DuckDB first, escalate only if you hit a wall." Five years ago that sentence would have been laughed out of an architecture review. Today it's the default.

The interop layer that made it inevitable

The other piece of the story is that DuckDB stopped being just an engine and became a federation point. It can read Parquet, CSV, JSON, Iceberg, Delta, Postgres, MySQL, SQLite, and Excel directly. It can write to all of those. It can run in the browser via WASM. It can be embedded in Python, R, Node, Java, Rust, Go, and a dozen other languages. The newer Vortex columnar format gives further demonstrated TPC-H gains over Parquet. It is the closest thing the data ecosystem has to a universal adapter.

That universality is what made the displacement irreversible. Once a tool can read everything, write to everything, and run anywhere — and once it does columnar scans at warehouse speed on a single machine — the question stops being "should I add a warehouse to this project" and starts being "do I have a real reason not to start with DuckDB."

For most projects in 2026, the answer is no.

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Data Engineering10 min read

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.

59 views
Read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Data Engineering19 min read

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.

58 views
Read
DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower
Data Engineering15 min read

DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower

DuckDB has no GPU backend of its own, and the GPU engines built for it need an NVIDIA card. gpudb 0.7 is my Apache-2.0 DuckDB extension for the GPU already inside your Mac, and for CUDA too. You write plain DuckDB SQL; the GPU takes a statement only where it has been measured faster than DuckDB on your own machine. On an Apple M4 Max, 19 of 22 TPC-H SF10 queries run on the Metal GPU at 1.06x to 48x with zero rows differing. The one row below parity is printed, not dropped.

78 views
Read