Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.
Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. I know, because I build one of these libraries, and I published the 50-million-row numbers too. So this time I measured the other thing: for 111 operations, the exact row count at which the GPU on an Apple M4 Max first stays ahead of the fastest CPU code I could find, and below which a single CPU core beats it. The answer is later than the benchmarks imply, sixteen operations never get there at all, and the fix I shipped is to stop using the GPU below the line.
Apple M4 Max, 16 CPU cores, 64 GB. Each operation timed at 1,000 to 50,000,000 rows against the fastest of Polars (eager and lazy), pyarrow (eager and 16-batch Acero), pandas and numpy; best of up to five after a warm-up, 10% nulls. From ArrowMetal's docs/CROSSOVER.md, generated 24 September 2026.
Everything here is from one Apple M4 Max with 16 CPU cores and 64 GB, and every number names its row count. The CPU side is not a strawman: it is the fastest of Polars (eager and lazy), pyarrow (eager and 16-batch Acero), pandas and numpy at each size, or, where I say so, a single-core loop. The generated page behind this post is docs/CROSSOVER.md in the ArrowMetal repository, an Apache-2.0 Arrow compute library for the Apple silicon GPU, listed on Arrow's Powered By page. Nothing on that page is typed by hand.
Why the GPU is slower than the CPU on small data: the 100-microsecond tax
Ask the GPU to sum 1,000 integers and it takes 112 microseconds. Polars does it in less than one. That is not a slow kernel. The kernel is done in a few nanoseconds. What you are paying for is the command-buffer round trip: encode, commit, wait, read back. On Apple silicon that floor is roughly 100 microseconds per call, and it does not care how many rows you sent.
| operation, 1,000 rows | ArrowMetal GPU | fastest CPU | which | GPU is behind by (at least) |
|---|---|---|---|---|
| sum(int64) | 0.112 ms | < 0.001 ms | Polars | > 100× |
| filter(int64 > 0) | 0.154 ms | 0.003 ms | Polars | > 50× |
| group-by sum, 1,000 keys | 0.401 ms | 0.080 ms | pandas | 5× |
Unified memory does not help here. People hear "no copy between CPU and GPU" and assume the GPU is free to call; it is not. Unified memory removes the transfer across a bus, which is the cost that makes medium-sized work pointless on a discrete card. The dispatch floor is a different cost, and the Mac pays it in full. What unified memory does is move the crossover much earlier than it sits on a PCIe GPU. It does not remove it. The way to remove it would be a persistent GPU kernel that polls shared memory for work instead of being dispatched; on Metal that cannot currently be made to work, for CPU/GPU coherence reasons the project measured and has filed with Apple as FB24858235. Batching many kernels into one command buffer shares the floor across them, which is how a chain of operations escapes it, but a single small call cannot.
Read the amber line. From 1,000 rows to 1,000,000 rows the GPU's time goes from 115 microseconds to 197. A thousand times more data, 70% more time. Meanwhile one CPU core goes from 1 microsecond to 993. They cross at 220,881 rows. Below that, using the GPU for a sum is a way of turning a 50-microsecond job into a 120-microsecond one.
GPU vs CPU: where the crossover actually is, for 111 operations
The single-core loop is the router's question, and I will come back to it. The question you are asking is harder: from what row count does the GPU stay ahead of the best CPU library, using all the cores it wants? Here is the answer by family. The bar runs from the operation that crosses earliest to the one that crosses latest; the ring is the median.
Two caveats on the number 111, so nobody has to find them for me. It is the set of operations the size sweep covers, not a selection: five more families, chains, decimal, join, temporal and window, are measured only at 10,000,000 and 50,000,000 rows in the full matrix and so have no crossover to state yet. And "median" above is the median of the 111 operations, which is 10,000,000 rows against the fastest multi-core library; 40 operations cross by 1,000,000, 55 need 10,000,000 or more. Three things that surprised me, in order of how much they should change what you do.
- Group-by is the early winner, not arithmetic. 27 of 28 group-by operations cross over, half of them by 100,000 rows. A sum by 100,000 keys is 1.4x ahead at 100,000 rows and 8.4x at 10,000,000. A GPU is a very good hash table.
- Element-wise arithmetic is the late one. Adding two int64 columns does not stay ahead of Polars until 10,000,000 rows. The work per element is one instruction and the CPU library runs it on 12 cores at memory bandwidth. There is nothing for the GPU to win except at sizes where its bandwidth advantage finally shows.
- Compare and select mostly never wins. Only 4 of 9. A filter that keeps 90% of rows is 0.87x at 50,000,000; a scalar compare is 0.79x. These are single passes at bandwidth on both sides and the GPU's dispatch floor never gets paid back.
The sixteen operations where the GPU never beats the CPU, and why
Sixteen of the 111 operations were still behind the fastest CPU code at the largest size measured. I would rather list them than let the chart imply otherwise. Every one has its cause in the project's to-improve list.
| operation | GPU ÷ fastest CPU at the largest size | why the GPU does not get there |
|---|---|---|
| ln, sin, sqrt, power (float64) | 0.10x, 0.16x, 0.87x, 0.95x | Metal has no double; float64 maths is software binary64, tens of integer instructions per element against one hardware instruction on 14 cores. |
| slice (zero-copy view) | 0.33x | The CPU library returns a view and moves nothing. ArrowMetal materialises. |
| unique, value_counts (1,000 distinct) | 0.83x, 0.84x | Answered by a full GPU radix sort plus a run scan; the CPU libraries hash. A hash kernel is the fix. |
| match_substring_regex (real pattern), parse, to_strings | 0.08x, 0.17x, 0.60x at 10M | Host fallbacks: a true regex runs NSRegularExpression on the CPU; number parsing and formatting are not GPU kernels yet. |
| drop_null, filter keeping 90%, replace_with_mask, case_when, compare scalar | 0.91x, 0.87x, 0.36x, 0.94x, 0.79x | Single-pass, memory-bound work where Polars lazy already runs near bandwidth on 10 to 14 cores; there is no 3x available to either side. |
| variance by key (1,000 groups) | 0.96x | Two-pass exact variance against a fused one-pass on the CPU. |
Two of those rows are the whole story of an earlier post: Apple GPUs have no 64-bit float hardware, so ln and sin on float64 are software arithmetic and stay at a tenth of Polars' speed. I would not change that decision; I would not use the GPU for logarithms either.
Where the GPU wins by 10x to 100x
The same machine, the same library, once the row count is big enough:
| operation, 50,000,000 rows | crossover | at 1M | at 10M | at 50M |
|---|---|---|---|---|
| any / all over booleans | 100,000 | 3.1x | 22.5x | 106x |
| quantile (median), float64 | 1,000,000 | 2.4x | 13.2x | 38.7x |
| partition_nth_indices | 1,000,000 | 2.3x | 17.5x | 24.5x |
| take, 25M random indices | 1,000,000 | 1.4x | 11.3x | 24.2x |
| lexsort, two int32 keys | 1,000,000 | 1.8x | 16.2x | 24.0x |
| list by key, 1,000 groups | 1,000,000 | 1.2x | 9.5x | 10.9x |
| argsort float64 | 1,000,000 | 4.7x | 10.0x | 10.5x |
| is_in, 100-value set | 10,000,000 | 0.57x | 1.8x | 7.4x |
Gathers, sorts, order statistics and anything that is secretly a sort. Those are the operations where a few thousand GPU lanes doing simple work beat a dozen cores doing clever work, and the gap widens as the data grows. If your job is 50,000,000 rows and a median, the GPU is 38.7x ahead of Polars and it is not a contest. If your job is 80,000 rows and a sum, the GPU is the wrong tool and the benchmark page was never going to tell you.
What I did about it: the library now refuses the GPU below the line
ArrowMetal 0.2.0, released 25 September 2026, ships a CPU/GPU router. For seven operations there is now a second implementation: a tight, single-threaded CPU loop over the Arrow layout. The router runs the loop below a measured crossover and the GPU kernel at or above it. The crossover table is not typed; it is generated by a script from a timing run of the shipped loops against the shipped kernels, and a test fails if the committed table stops matching the results file it names.
| routed operation | CPU loop wins below | the two sizes it was fitted between (GPU µs / CPU µs) |
|---|---|---|
| sum | 220,881 rows | 100K: 125.6 / 46 · 300K: 161.5 / 213.6 |
| max | 227,546 | 100K: 126.3 / 66.8 · 300K: 178.4 / 212.2 |
| min | 234,408 | 100K: 124.2 / 65.8 · 300K: 180.3 / 208.8 |
| filter by mask | 591,365 | 300K: 375 / 198.1 · 1M: 406.8 / 654.9 |
| multiply | 678,287 | fitted from the multiply loop; a 64-bit integer multiply costs the CPU more per element than an add |
| group-by sum, ≤ 1,024 keys | 716,029 | 300K: 269.9 / 186.1 · 1M: 558 / 615.2 |
| add, subtract | 1,629,041 | 1M: 181.8 / 120.5 · 3M: 230.1 / 363.7 |
| compare | 2,668,823 | 1M: 141.5 / 77 · 3M: 211.2 / 224 |
Three rules I held myself to, because a router that returns slightly different answers on each path would be worse than no router:
- Byte-identical output. Same values, same validity bitmap, same null count, same buffer sizes. Integer arithmetic wraps as the GPU wraps. Float sums do not add in row order: the CPU loop reproduces the GPU's summation order exactly, thread by thread and threadgroup by threadgroup, so the sum is the GPU's to the bit, NaN payloads included. The tests hold both paths to that across sizes, null densities, offsets, signed zeros and signalling NaNs.
- No CPU path, no choice. Float32 arithmetic computes in hardware
floaton the GPU, where subnormals are flushed, so a CPU loop could not match its bits; it stays on the GPU and the decision says why. So does anything inside a batch, since batching exists to share one dispatch floor across many kernels. - Measured, then checked. Re-running the check against the fitted table,
autotook the faster path in 71 of 72 cases. The one miss iscompare(int64 > int64)at 3,000,000 rows, by a margin of 1.10, because the compare row was fitted on the scalar form.
Strings are not routed, and the reason is worth stating: the string rows are behind a 12-core Acero, and a single-threaded loop would not change that. The router chooses between ArrowMetal's own two paths. It never pretends to be Polars, and it means a small string operation still pays the dispatch floor today.
You can watch it decide. This ran on the M4 Max against the 0.2.0 wheel from PyPI; the comments are the actual output:
pip install -U arrowmetal # 0.2.0 · macOS 14+ on Apple silicon
import numpy as np, pyarrow as pa, arrowmetal as am
am.device_name() # 'Apple M4 Max'
am.router_crossovers() # {'sum': 220881, 'min': 234408, 'max': 227546, 'compare': 2668823,
# 'arithmetic': 1629041, 'filter': 591365, 'group_by_sum': 716029, 'multiply': 678287}
for n in (1_000, 100_000, 1_000_000, 10_000_000):
col = am.array(pa.array(np.arange(n, dtype=np.int64)))
col.sum()
r = am.last_route()
print(f"{n:>11,} rows -> {r.path} ({r.reason})")
# 1,000 rows -> cpu (below the 220881-row crossover)
# 100,000 rows -> cpu (below the 220881-row crossover)
# 1,000,000 rows -> gpu (at or above the 220881-row crossover)
# 10,000,000 rows -> gpu (at or above the 220881-row crossover)
with am.router("gpu"): # or "cpu"; or ARROWMETAL_ROUTER=gpu|cpu|auto for the process
col.sum()
am.last_route().reason # 'forced gpu by the per-thread override'
The rest of 0.2.0 is the bigger half
The router is the part of this release that should change how you think. It is not the part that took the time. 0.1.0 was a compute library: 307 Arrow kernels that were very fast once your columns were already on the GPU, and no opinion about how they got there. 0.2.0 is the version where the data arrives.
| 0.1.0, 8 September | 0.2.0, 25 September |
|---|---|
| 307 Arrow compute kernels on the GPU, always the GPU | A CPU/GPU router with a measured crossover table; seven operations run a single-core loop below it, byte-identical |
| Parquet: flat columns | Parquet on the GPU: structs, maps, lists nested to any depth, the stored Arrow schema applied, page-index and bloom-filter skipping |
| No text readers | GPU readers for CSV and newline-delimited JSON, checked against pyarrow's |
| Files only | Delta Lake and Apache Iceberg tables read through the GPU Parquet reader |
| Polars: convert a Series over the C Data interface | lf.collect(engine=am.MetalEngine()): subtrees of the optimised Polars plan replaced with GPU plans, the rest left to Polars |
| DuckDB: a Python bridge and a loadable extension | A DuckDB optimizer extension that moves eligible aggregates of unchanged SQL onto the GPU |
| Arrow IPC: the common types | Every Arrow type nested recursively, LZ4 and ZSTD bodies, big-endian files, the view types, the fixed-shape tensor extension |
| String sort: nulls partitioned on the CPU | Nulls carried in the GPU sort key: 10,000,000 strings with 10% nulls, 122.7 ms to 42.0 ms |
Parquet, nested, with the skipping a lakehouse actually needs
A real Parquet file is not a table of integers. It is structs of structs of strings, maps with nested values, lists inside lists inside structs. 0.2.0 reassembles all of it on the GPU from the leaf columns: three levels per field from the schema, then one flag kernel, one prefix sum and one scatter per field. It is checked value for value against pyarrow.parquet.read_table on files written by pyarrow, DuckDB and Polars, because those three do not always agree on how to write the same thing.
The part that matters for a query is what it does not read. With a column index in the file, pages whose min and max cannot match the filter are never read, decompressed or decoded; a row group every page of which is ruled out is dropped; and every column is trimmed to the same candidate rows. Split-block bloom filters, the kind pyarrow and DuckDB write, drop row groups for an equality filter before any page is touched. The matching rows are identical with and without the index, and last_read_stats tells you how many pages were decoded and how many skipped. There is also a bug-fix in here I am glad was found before you did: a list column with more than 4,096 level entries in one read used to write past a 4-byte scratch buffer into host memory. It has a slot per level now, and a test that fails without it.
CSV, JSON, Delta Lake and Iceberg
CSV and newline-delimited JSON now parse on the GPU, checked against pyarrow's readers. Delta Lake and Apache Iceberg tables are read through the GPU Parquet reader, which is the first time I know of that either format has been opened on an Apple GPU at all. I will say the next part plainly rather than let you find it after a migration: on day one both lakehouse readers are behind the fastest CPU reader, and so are the IPC view layouts and some nested Parquet reads. The measurements are in Benchmarks/results/, the roadmap lists them, and the lakehouse page says which. The crossover discipline from the first half of this post applies to readers too; I just have not finished measuring it for them.
Polars gets a GPU engine
Polars has a hook that hands its optimised query plan to an engine after optimisation. 0.2.0 plugs into it: lf.collect(engine=am.MetalEngine()) walks every node and expression of the plan, translates the subtrees it can into an ArrowMetal plan, and replaces each with a function that runs on the GPU and returns a Polars DataFrame. Everything it cannot take stays with Polars, which also runs whatever sits above a replaced subtree. The most that can go wrong is that nothing moves and you get plain Polars' answer. engine.last_report tells you which nodes ran on Metal and why the rest did not.
Building that engine is also where four wrong-answer bugs came from, because it runs a differential suite of Polars queries against the GPU plans and Polars exports things the hand-written tests never did. The best of them: a string filter or take that met a null slot which still held bytes (valid Arrow; Polars exports them) copied those bytes over the next kept row, and turned "banana" into "xanana". It is fixed, it has a test, and the findings log lists the other three.
DuckDB rewrites your SQL, or leaves it alone
The DuckDB side gained an optimizer extension: loaded into a connection, it moves eligible aggregates of ordinary queries onto the GPU with the SQL unchanged, and SET arrowmetal_rewrite = 'off' or 'force' turns it off or overrides the decision. It now refuses an unrecognised setting value with an error naming the accepted ones, where it used to silently treat it as 'auto'. If that sounds like the never-slower rule from gpudb, it is the same instinct arriving from the other direction: gpudb is a GPU under DuckDB, ArrowMetal is Arrow compute that DuckDB and Polars can both borrow.
And the string sort got three times faster by moving a wait
The utf8 sort used to partition its null rows on the CPU, which meant waiting for the GPU batch that produced the sort keys. 0.2.0 carries the null placement in bit 63 of every prefix key, a bit the seven 9-bit byte fields leave free, so the stable radix passes leave the nulls as one block at either end and the CPU wait is gone. A plan sort of 10,000,000 strings of 1 to 24 bytes with 10% nulls went from 122.7 ms to 42.0 ms on the M4 Max. Columns without nulls take the same passes as before. That change also closed a real bug: with two or more prefix passes the null partition could read the index array before the GPU had written it, which returned wrong rows and, once, a bus error.
The rule I would actually give you
Do not ask whether the GPU is faster. Ask at how many rows, for this operation, on this machine, and whether your data is already in Arrow memory so the crossing costs nothing. On an Apple M4 Max the honest answers are: group-by and strings from about 100,000 rows; sorts, gathers and order statistics from about 1,000,000; element-wise arithmetic from about 10,000,000; and never for float64 transcendentals, views, or anything a multi-core CPU already runs at bandwidth. Below those lines a single core is the right tool, and as of 0.2.0 the library picks it for you, then reads your Parquet, your Delta table and your Polars plan on the other side of the line.
ArrowMetal is an independent Apache-2.0 project that implements Apache Arrow, listed on Arrow's Powered By page; it is version 0.2.0 and every number above is from one machine. pip install -U arrowmetal on macOS 14 or later, arrowmetal = "0.2.0" on crates.io. If you run Benchmarks/router_check.py on a different Apple silicon Mac, the crossovers you get are new information and I would like to see them as an issue on GitHub.
FAQ: GPU vs CPU for data work
Is a GPU faster than a CPU for data processing?
Only above a row count that depends on the operation. Measured on an Apple M4 Max against the fastest of Polars, pyarrow, pandas and numpy, the median of 111 operations first stays ahead on the GPU at 10,000,000 rows (40 cross by 1,000,000; 55 need 10,000,000 or more); against a single CPU core the line is nearer 1,000,000. Group-by and string operations cross earliest, from 100,000; element-wise arithmetic latest; and 16 of the 111 never overtake the CPU at any size measured, up to 50,000,000 rows.
At how many rows does a GPU become faster than a CPU?
For the operations ArrowMetal 0.2.0 routes, the crossover against a single-core CPU loop on an M4 Max is 220,881 rows for sum, about 230,000 for min and max, 591,365 for filter, 716,029 for a 1,024-key group-by sum, 1,629,041 for add and 2,668,823 for compare. Against a multi-core CPU library the line moves later; the crossover document in the repository gives every operation.
Why is the GPU so slow on small data?
Every GPU call pays a fixed dispatch cost, a command-buffer round trip, of roughly 100 microseconds on Apple silicon whatever the input size. Summing 1,000 integers takes the GPU 112 microseconds and Polars under one. Below a few hundred thousand rows that floor is the whole cost, and unified memory does not remove it; it removes the copy, which is a different cost.
Which operations should never go on a GPU?
From the 111 measured: float64 transcendentals such as ln and sin (no hardware double on Apple GPUs), zero-copy views such as slice, hash-style operations answered by a sort such as unique and value_counts, real regular expressions and number parsing (host fallbacks), and single-pass memory-bound operations where a multi-core CPU already runs at bandwidth. Each has its cause written down in the project's to-improve list.
Does Apple silicon unified memory make the GPU faster for small data?
No. Unified memory removes the copy across a bus, which is what makes medium-sized jobs worth moving at all on a Mac. It does not remove the dispatch floor of about 100 microseconds per call, so the crossover still exists; it is simply much lower than on a discrete GPU where the transfer is added on top.
What does ArrowMetal 0.2.0 do below the crossover?
For sum, min, max, mean, compare, add, subtract, multiply, filter and small-key group-by sums it runs a single-threaded CPU loop instead of the GPU kernel, chosen from a measured crossover table, and returns a byte-identical Arrow array, including float sums, which reproduce the GPU's summation order bit for bit. A check against the shipped table took the faster path in 71 of 72 cases.
Read next
- Apple's GPU has no 64-bit floats. I made it sort 50 million doubles anyway
- Polars vs DuckDB vs ArrowMetal GPU on Apple silicon: sort and group-by benchmarks
- DuckDB on the Apple silicon GPU: plain SQL, and a rule that says never slower
- ArrowMetal: Apache Arrow compute on the Apple silicon GPU, the launch post
References & Citations
- ArrowMetal (2026). “The crossover: where the GPU path overtakes the CPU.” docs/CROSSOVER.md, generated 24 September 2026 by Benchmarks/crossover.py on an Apple M4 Max, 16 cores, 64 GB; 111 operations in six families against Polars 1.44.1 (eager and lazy), pyarrow 25.0.1 (eager and Acero), pandas 3.0.5 and numpy 2.5.3. Source of the family chart, the never-win table and the wins table.
- ArrowMetal (2026). Sources/ArrowMetal/Router/RouterTable.swift, generated from Benchmarks/results/router_check_2026-09-24.csv (int64, 10% nulls, best of 20). Source of the router table and the sum crossover chart.
- ArrowMetal (2026). docs/DESIGN.md, “CPU/GPU router”: what is routed, the decision order, byte-identical output, the 71-of-72 check.
- ArrowMetal (2026). docs/BENCHMARKS_MATRIX.md, latency family, generated 7 September 2026. Source of the 1,000-row floor table. docs/TO_IMPROVE.md for the causes.
- ArrowMetal (2026). Release v0.2.0, 25 September 2026: readers, Parquet, Delta Lake and Iceberg, the Polars engine, the DuckDB optimizer extension, the string-sort figure and the findings.
- The Python snippet was run on 24 September 2026 on an Apple M4 Max, Python 3.13.9, arrowmetal 0.2.0, pyarrow 25.0.1 and numpy 2.5.3 from PyPI; the comments are its actual output.
Subscribe to new posts from theaivibe.org
Related Posts

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.
DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower
DuckDB has no GPU backend of its own, and the GPU engines built for it need an NVIDIA card. gpudb 0.7 is my Apache-2.0 DuckDB extension for the GPU already inside your Mac, and for CUDA too. You write plain DuckDB SQL; the GPU takes a statement only where it has been measured faster than DuckDB on your own machine. On an Apple M4 Max, 19 of 22 TPC-H SF10 queries run on the Metal GPU at 1.06x to 48x with zero rows differing. The one row below parity is printed, not dropped.
Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway
The Metal Shading Language has no double type, and float64 is the default number in Python, pandas and Apache Arrow. Building ArrowMetal meant getting past three walls: a GPU with no 64-bit floats, a missing 64-bit atomic add, and a Swift compiler bug that reports errors nobody threw. Here is how each one was solved, what it cost, and why a GPU that cannot add two doubles sorts 50,000,000 of them in 32 ms.