Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still
Polars 2.0.0 shipped on 6 October and runs its group-bys, joins and sorts faster than 1.44. Two days later ArrowMetal's MetalEngine was re-measured against it on an Apple M4 Max: the GPU default takes 52 of 107 queries at 50 million rows, each 1.35x to 7.94x faster than the faster Polars 2.0 engine, and hands 12 back that it used to take. What Polars 2.0 changed, and how a crossover table answers it.
Polars 2.0.0 came out on 6 October, and it is faster. On my Apple M4 Max, over the 62 group-by, join, sort and unique queries that ArrowMetal's GPU engine used to take from Polars 1.44.1 at 50,000,000 rows, Polars 2.0.0's own time is 0.72x of 1.44.1's at the median; unique over two int32 keys went from 153.1 to 38.5 ms. That is the kind of release that makes a GPU engine's benchmark table wrong, so within 48 hours ArrowMetal was re-measured against it, case by case, and its default refitted. This is what came out.
Apple M4 Max, Polars 2.0.0 (16 threads), ArrowMetal main after 0.4.0; 50,000,000 rows in memory, cold, best of 5, AC power, 8 October 2026.
ArrowMetal is my Apache-2.0 Arrow compute library for the Apple silicon GPU, on Arrow's Powered By page. Its MetalEngine is a Polars engine: lf.collect(engine=engine) returns the frame lf.collect() returns, the subtrees it was measured ahead on run on the GPU, the rest by Polars. The DataFusion side shipped last week; this is the Polars side.
What Polars 2.0 changed, as a benchmark sees it
The release notes are long; the benchmark's view of them is short. The streaming engine is now the default and the faster one for 56 of those 62 queries, group-bys and joins got quicker, and the GPU's own times did not move, a median ratio of 1.03 between the runs. Two queries went the other way, and two things the engine depended on were removed.
| what Polars 2.0.0 changed | as the engine and the benchmark saw it, Apple M4 Max, 50,000,000 rows |
|---|---|
collect() runs the streaming engine | On 1.44.1 a plain collect() ran the in-memory engine. Over the 62 queries the 1.44 default took at 50,000,000 rows, 2.0.0's faster engine is the streaming one in 56. |
| Group-bys, joins and sorts are faster | 0.72x of 1.44.1's time at the median of those 62; unique over two int32 keys 153.1 to 38.5 ms. The GPU's times on the same queries: median ratio 1.03. |
| Two queries got slower | Filter-then-sort over int32 and int64 keys, 694 to 713 ms; unique over a String and an int32 key, 256 to 296 ms. |
LazyFrame.profile and with_context are gone | The engine's profile raises NotImplementedError on 2.0.0 and says why; the plan walker has 19 node kinds to know, not 20. |
| The IR moved from (14, 7) to (15, 2) | One tested IR version per major; an unseen major keeps every plan with Polars. The 2,523-plan capability grid gives the same 1,326 to Metal on both; 63 plans 1.44.1 ran, 2.0.0 itself rejects. |
The engine ran on 2.0.0 the same evening, its tests passing. The open question was whether the default, the shapes and row counts it takes without being asked, was still true. Eight of the 107 queries tell most of it:
Read the amber labels. On the two-key unique the GPU was 11.54x ahead of Polars 1.44.1 and is 2.83x ahead of 2.0.0, at the same 13 ms; Polars moved. The int64 sort went from 7.17x to 2.82x the same way. On the filter-then-sort and the String unique, 2.0.0 is a little slower than 1.44.1 and the GPU's lead grew to 7.94x and 7.03x. The one case under 1x, the sort with a String column at 0.89x, is why the default needed a new table rather than a footnote.
One crossover table per Polars major
The default is a measured table, not a flag. For each shape the engine can translate, a sweep runs it from 250,000 to 50,000,000 rows through both Polars engines and the GPU, cold, best of 7, and the fit finds the first size from which the GPU stays ahead by a margin, 15% for numeric shapes and 35% with a String column, at every larger size. The default takes the shape from 1.5 times that fit, and not at all if that passes 50,000,000 rows. Group-bys add the number of groups, estimated at run time from a 512- or 2,048-row sample with the Chao1 estimator, because the same count is ahead at 1,000,000 groups and behind at 200.
There is now one such table per Polars major: MetalEngine() loads the 1.44.1 table on Polars 1.x and the 2.0.0 table on 2.x, the first time the policy decides. Where the rows differ:
| shape | where the default starts taking it: Polars 1.x table against Polars 2.x table |
|---|---|
| Numeric full sort | from 1,000,000 rows on 1.x to 10,000,000 on 2.x. |
unique over numeric keys | from 5,494,090 rows to 2,943,034: Polars 2.0.0 is 4x faster here, yet the GPU's lead now starts earlier, because the 1.x fit was held up by a case 2.0.0 no longer runs slowly. |
| Inner and anti join | inner from 2,029,827 rows to 1,942,068; anti from 3,727,959 to 1,875,000. Left join unchanged at 1,875,000. |
| Sort with a String column | taken from 5,000,000 rows on 1.x; not taken on 2.x. At 50,000,000 rows the GPU is 0.89x, 244.77 ms against 217.88, so its class is left. |
Group-by sum and mean at 200 to 10,000 groups | taken in several buckets on 1.x; left on 2.x except a two-key sum at 1,000 groups from 28,976,995 rows and a one-key mean at 10,000 groups from 50,000,000. |
The sort row is the one to notice. The GPU sorts 50,000,000 int64 keys in the same 52.70 ms on both Polars versions; what changed is that 2.0.0's streaming sort at 5,000,000 rows takes 13.55 ms where 1.44.1's took 32.88, so the size from which the GPU is reliably ahead moved from one million rows to ten.
The 12 queries handed back
With the 2.x table in force, the default takes a subtree in 52 of the 107 in-memory queries at 50,000,000 rows, every result equal to Polars', and is ahead of the faster Polars 2.0.0 engine in all 52, from 1.35x on an inner join followed by a whole-frame sum to 7.94x on the filter-then-sort. Under the 1.x table it took 63, one behind. The difference:
| handed back to Polars 2.0.0 | what the run measured, and why the table leaves it |
|---|---|
Ten group-by sum and mean queries, 200 to 10,000 groups | 1.15x to 1.88x ahead at 50,000,000 rows in this run, but 0.36x to 0.97x at 250,000 and 500,000 rows in the sweep; the fit wants a bucket ahead by 15% from its crossover up, with 1.5x of headroom, and no such row exists for these. |
| Sort with a String column, by an int64 key | Behind: 0.89x at 50,000,000 rows, 244.77 ms against Polars 2.0.0's 217.88. To improve. |
| Filter, then sort by (int32 asc, nullable Float64 desc) | Ahead, 2.75x, 317.71 ms against 874.34. Left anyway: it shares a class with the String-column sort, and a class is taken only when every case measuring it is ahead. |
| One query newly taken | Group-by count over one key at 1,000 groups, 1.91x at 50,000,000 rows, the 2.x table's one addition. |
The ten group-bys are the honest part. Nine were ahead in this run, some by 1.8x, and the default still leaves them, because the rule is not "ahead at 50,000,000 rows" but "ahead from a crossover, with headroom", and at 250,000 and 500,000 rows Polars 2.0.0 does a two-key sum over 1,000 groups in a few milliseconds the GPU cannot match. MetalEngine(shapes="all") takes every translatable subtree of a million rows or more; either way the report names the rule behind each node. From my machine on 8 October, Polars 2.0.0, a 50,000,000-row frame with a permuted int64 key:
This is what the engine printed on my machine on 8 October, Polars 2.0.0, a 50,000,000-row frame with a permuted int64 key:
import polars as pl, arrowmetal as am # polars 2.0.0
engine = am.MetalEngine() # the measured defaults, the 2.x table
df = lf.sort("id").collect(engine=engine) # the frame lf.collect() returns
print(engine.last_report)
MetalEngine report for collect (polars 2.0.0, IR (15, 2))
metal: Sort#1 [Sort > DataFrameScan] over 50,000,000 rows, ran in 73.09 ms -> 50,000,000 rows
rule: 50,000,000 input rows is at or above the 10,000,000-row crossover for sort
(engine table, Benchmarks/results/polars_engine_crossover_2026-10-08-polars2.csv)
# the same plan: Polars 2.0.0 collect() 199.74 ms (streaming), engine="in-memory" 410.51 ms;
# df.equals(lf.collect()) -> True
And the two kinds of no:
lf.group_by("region").agg(pl.col("amount").mean()).collect(engine=engine)
MetalEngine report for collect (polars 2.0.0, IR (15, 2))
nothing ran on Metal
polars: GroupBy#1: rule: estimated 50 groups over (region), a Chao1 estimate from a
512-row sample, 50 to 50: below the measured band for group_by:mean at
50,000,000 input rows (taken at 3,163 to 3,162,277 groups; ...)
groups: GroupBy#1 over (region): 50 groups ... (probed in 94 us)
engine.profile(lf)
NotImplementedError: MetalEngine.profile needs LazyFrame.profile, which polars 2.0.0
does not have (Polars 2.0 removed it).
The answers, to the bit
Every one of the 52 taken queries in the refit run is marked equal to Polars'. Counts, keys, join rows and sort orders match exactly; Polars sorts floats in a total order with its null placement per key, and the GPU sort takes both as options on its one radix pass. Float64 group sums are the one place a GPU and a CPU can legitimately differ, in the last bit, because they add in a different order. On a two-key sum of 50,000,000 uniform Float64 values into 100,000 groups the two agree exactly on 11,917 groups and within 2.2e-15 relative on the rest; on the one group I checked against an exactly rounded reference, the GPU's sum is the correctly rounded one.
Install, honestly
pip install "arrowmetal[polars]"today installs 0.4.0, which pins Polars to 1.44. The 2.0.0 support and the second table are onmainas of 8 October, in the CHANGELOG under "Unreleased", and go out in the next release; until then it is a source build, three commands in docs/POLARS.md, and the extra becomespolars>=1.44,<2.1.MetalEngine.profileraises on 2.0.0; the rest of the API is the same on both majors, andpython -m arrowmetal.polars_engine checknames the Polars, the IR and the table in force.- macOS 14 or later on Apple silicon; every number here is one M4 Max on AC power, load average logged.
python -m arrowmetal.benchmeasures yours in under 30 s.
ArrowMetal is an independent Apache-2.0 project that implements Apache Arrow; Apache Arrow is a trademark of the Apache Software Foundation, and Polars is its maintainers' project, of which I am a user. The table gets refitted again when 2.1 moves the numbers.
FAQ: Polars 2.0, the GPU and Apple silicon
Does Polars 2.0 run on the Apple silicon GPU?
With ArrowMetal's MetalEngine, the subtrees it was measured ahead on do: full sorts, unique, joins and group-bys judged by their group count; Polars runs the rest. The 2.0.0 support is on main, in the next release.
Is Polars 2.0 faster than Polars 1.44?
On an Apple M4 Max at 50,000,000 rows, yes: over 62 group-by, join, sort and unique queries its time is 0.72x of 1.44.1's at the median, 0.25x on unique over two keys. collect() now runs the streaming engine.
How much faster is the GPU than Polars 2.0?
The default takes 52 of 107 benchmark queries at 50,000,000 rows and is 1.35x to 7.94x faster than the faster Polars 2.0.0 engine in every one, none behind.
Why does the GPU engine hand 12 queries back on Polars 2.0?
Polars 2.0.0 got faster on them, so the crossover table has no row where the GPU is ahead by its 15% margin with headroom. The default is fitted per Polars major and takes only what it measured ahead.
Read next
References & Citations
- Polars (2026). Python Polars 2.0.0, GitHub release py-2.0.0, 6 October 2026; PyPI upload 2026-10-06 11:44 UTC.
- ArrowMetal (2026). “ArrowMetal for Polars users.” docs/POLARS.md: “Which Polars”, “How it plugs into Polars 1.44.1 and 2.0.0”, “Which translatable subtrees it runs: the defaults” incl. “Four rules on the fit” and “On polars 2.0.0”; the README's “Polars in full”; the CHANGELOG, “Unreleased”, as of main c15d992 (8 October 2026).
- ArrowMetal (2026). Benchmarks/results/polars_engine_bench_2026-10-08-polars2-refit.csv: the 52 of 107, the 1.35x to 7.94x, the chart's 2.0.0 and MetalEngine bars. polars_engine_bench_2026-10-02.csv: the chart's 1.44.1 bars and the 75-of-220 baseline. polars_engine_bench_2026-10-07-polars2.csv: the 0.72x median and the 56 of 62. polars_engine_bench_2026-10-08-polars2-after.csv: the 63 under the 1.x table, the 12 handed back, the String sort at 0.89x.
- ArrowMetal (2026). polars_engine_crossover_2026-10-08-polars2.csv and python/arrowmetal/_engine_crossovers_pl2.py: the 2.x table and its fit. docs/ENGINE_CAPABILITIES_POLARS2.md: 1,326 of 2,523 plans, 63 n/a.
- Chao, A. et al. (2005). The bias-corrected Chao1 estimator, Biometrics, as cited in docs/POLARS.md.
- The two code cards and the Float64 group-sum check are from a run for this post on 8 October 2026 (Polars 2.0.0, ArrowMetal main after 0.4.0 built from source, Apple M4 Max, macOS 27.0, 50,000,000 rows in memory); the report lines are the engine's actual output, trimmed for width.
Subscribe to new posts from theaivibe.org
Related Posts
DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster
Apache DataFusion runs any physical optimizer rule you register. ArrowMetal 0.4.0 ships one for the Apple silicon GPU: the SQL is unchanged, the answers are DataFusion's, and full sorts of 250,000 to 50 million rows run 6.9x to 28.8x faster on an M4 Max. What it takes, what it leaves, and why.

sqljev: TypeSafe Jev's jev() for SQL Server, Postgres, Snowflake, BigQuery and DuckDB, on Jev or Open-Weight Laya
SQL cannot say 'the customer threatens to cancel'. sqljev adds jev(), jev_prob() and jev_choice() to SQL Server, PostgreSQL, MySQL, Snowflake, Databricks, BigQuery, Redshift and DuckDB, and answers them with Laya, an open-weight decision model that runs on your hardware, returns calibrated probabilities instead of text, and can be fine-tuned on your own tables. 140,000 decisions in 271 seconds on one laptop GPU; re-running all 13 queries, 0.9 seconds. Apache 2.0, version 0.1.0.

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.