Back to Blog

Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still

Prateek SinghOctober 8, 202610 min read3 views
Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still

Polars 2.0.0 shipped on 6 October and runs its group-bys, joins and sorts faster than 1.44. Two days later ArrowMetal's MetalEngine was re-measured against it on an Apple M4 Max: the GPU default takes 52 of 107 queries at 50 million rows, each 1.35x to 7.94x faster than the faster Polars 2.0 engine, and hands 12 back that it used to take. What Polars 2.0 changed, and how a crossover table answers it.

Polars 2.0.0 came out on 6 October, and it is faster. On my Apple M4 Max, over the 62 group-by, join, sort and unique queries that ArrowMetal's GPU engine used to take from Polars 1.44.1 at 50,000,000 rows, Polars 2.0.0's own time is 0.72x of 1.44.1's at the median; unique over two int32 keys went from 153.1 to 38.5 ms. That is the kind of release that makes a GPU engine's benchmark table wrong, so within 48 hours ArrowMetal was re-measured against it, case by case, and its default refitted. This is what came out.

queries the GPU default takes on Polars 2.0.0
52 of 107
at 50,000,000 rows; every answer equal to Polars'
ahead of the faster Polars 2.0.0 engine
1.35x to 7.94x
in all 52; none behind
Polars 2.0.0's own time against 1.44.1
0.72x
at the median over 62 queries; 0.25x on unique over two keys
queries handed back to Polars 2.0
12
taken on 1.44; left now, 0.89x to 2.75x

Apple M4 Max, Polars 2.0.0 (16 threads), ArrowMetal main after 0.4.0; 50,000,000 rows in memory, cold, best of 5, AC power, 8 October 2026.

ArrowMetal is my Apache-2.0 Arrow compute library for the Apple silicon GPU, on Arrow's Powered By page. Its MetalEngine is a Polars engine: lf.collect(engine=engine) returns the frame lf.collect() returns, the subtrees it was measured ahead on run on the GPU, the rest by Polars. The DataFusion side shipped last week; this is the Polars side.

What Polars 2.0 changed, as a benchmark sees it

The release notes are long; the benchmark's view of them is short. The streaming engine is now the default and the faster one for 56 of those 62 queries, group-bys and joins got quicker, and the GPU's own times did not move, a median ratio of 1.03 between the runs. Two queries went the other way, and two things the engine depended on were removed.

what Polars 2.0.0 changedas the engine and the benchmark saw it, Apple M4 Max, 50,000,000 rows
collect() runs the streaming engineOn 1.44.1 a plain collect() ran the in-memory engine. Over the 62 queries the 1.44 default took at 50,000,000 rows, 2.0.0's faster engine is the streaming one in 56.
Group-bys, joins and sorts are faster0.72x of 1.44.1's time at the median of those 62; unique over two int32 keys 153.1 to 38.5 ms. The GPU's times on the same queries: median ratio 1.03.
Two queries got slowerFilter-then-sort over int32 and int64 keys, 694 to 713 ms; unique over a String and an int32 key, 256 to 296 ms.
LazyFrame.profile and with_context are goneThe engine's profile raises NotImplementedError on 2.0.0 and says why; the plan walker has 19 node kinds to know, not 20.
The IR moved from (14, 7) to (15, 2)One tested IR version per major; an unseen major keeps every plan with Polars. The 2,523-plan capability grid gives the same 1,326 to Metal on both; 63 plans 1.44.1 ran, 2.0.0 itself rejects.
From the README's "Polars in full", docs/POLARS.md "On polars 2.0.0" and the CHANGELOG. The 62-query comparison is the 2 October run on 1.44.1 against the 7 October run on 2.0.0, same code, same machine.

The engine ran on 2.0.0 the same evening, its tests passing. The open question was whether the default, the shapes and row counts it takes without being asked, was still true. Eight of the 107 queries tell most of it:

The same eight queries on Polars 1.44.1, Polars 2.0.0, and the GPU under Polars 2.0.0 50,000,000 rows in memory · Apple M4 Max · best of 5 · ms on a log scale · the amber label is Polars 2.0.0 ÷ MetalEngine Polars 1.44.1, faster engine Polars 2.0.0, faster engine MetalEngine() on 2.0.0 5 ms10 ms20 ms50 ms100 ms200 ms500 ms1,000 ms filter, then sort by (int32 asc, int64 desc) 694.1 713.2 89.82 · 7.94x unique over (String, int32), keep first 255.8 296.3 42.14 · 7.03x sort by a nullable Float64 key, descending 553.0 173.8 93.54 · 1.86x sort 3 columns by an int64 key 380.5 148.8 52.70 · 2.82x unique over 2 int32 keys, keep first 153.1 38.01 13.42 · 2.83x inner join, 1,000,000-row build side 82.91 57.97 22.63 · 2.56x group-by 1 key, 200 groups, min + max 16.46 17.20 10.54 · 1.63x inner join, then a whole-frame sum 15.49 8.24 6.12 · 1.35x
Eight of the 107 benchmark queries, 50,000,000 rows in memory on an Apple M4 Max, best of 5. Polars bars are the faster of its two engines on that version: 1.44.1 from polars_engine_bench_2026-10-02.csv, 2.0.0 and the MetalEngine from polars_engine_bench_2026-10-08-polars2-refit.csv.

Read the amber labels. On the two-key unique the GPU was 11.54x ahead of Polars 1.44.1 and is 2.83x ahead of 2.0.0, at the same 13 ms; Polars moved. The int64 sort went from 7.17x to 2.82x the same way. On the filter-then-sort and the String unique, 2.0.0 is a little slower than 1.44.1 and the GPU's lead grew to 7.94x and 7.03x. The one case under 1x, the sort with a String column at 0.89x, is why the default needed a new table rather than a footnote.

One crossover table per Polars major

The default is a measured table, not a flag. For each shape the engine can translate, a sweep runs it from 250,000 to 50,000,000 rows through both Polars engines and the GPU, cold, best of 7, and the fit finds the first size from which the GPU stays ahead by a margin, 15% for numeric shapes and 35% with a String column, at every larger size. The default takes the shape from 1.5 times that fit, and not at all if that passes 50,000,000 rows. Group-bys add the number of groups, estimated at run time from a 512- or 2,048-row sample with the Chao1 estimator, because the same count is ahead at 1,000,000 groups and behind at 200.

There is now one such table per Polars major: MetalEngine() loads the 1.44.1 table on Polars 1.x and the 2.0.0 table on 2.x, the first time the policy decides. Where the rows differ:

shapewhere the default starts taking it: Polars 1.x table against Polars 2.x table
Numeric full sortfrom 1,000,000 rows on 1.x to 10,000,000 on 2.x.
unique over numeric keysfrom 5,494,090 rows to 2,943,034: Polars 2.0.0 is 4x faster here, yet the GPU's lead now starts earlier, because the 1.x fit was held up by a case 2.0.0 no longer runs slowly.
Inner and anti joininner from 2,029,827 rows to 1,942,068; anti from 3,727,959 to 1,875,000. Left join unchanged at 1,875,000.
Sort with a String columntaken from 5,000,000 rows on 1.x; not taken on 2.x. At 50,000,000 rows the GPU is 0.89x, 244.77 ms against 217.88, so its class is left.
Group-by sum and mean at 200 to 10,000 groupstaken in several buckets on 1.x; left on 2.x except a two-key sum at 1,000 groups from 28,976,995 rows and a one-key mean at 10,000 groups from 50,000,000.
From docs/POLARS.md, "Where its rows differ from the 1.x table's", and _engine_crossovers_pl2.py, fitted from the 8 October sweep: 250,000 to 50,000,000 rows, best of 7, one process per size.

The sort row is the one to notice. The GPU sorts 50,000,000 int64 keys in the same 52.70 ms on both Polars versions; what changed is that 2.0.0's streaming sort at 5,000,000 rows takes 13.55 ms where 1.44.1's took 32.88, so the size from which the GPU is reliably ahead moved from one million rows to ten.

The 12 queries handed back

With the 2.x table in force, the default takes a subtree in 52 of the 107 in-memory queries at 50,000,000 rows, every result equal to Polars', and is ahead of the faster Polars 2.0.0 engine in all 52, from 1.35x on an inner join followed by a whole-frame sum to 7.94x on the filter-then-sort. Under the 1.x table it took 63, one behind. The difference:

handed back to Polars 2.0.0what the run measured, and why the table leaves it
Ten group-by sum and mean queries, 200 to 10,000 groups1.15x to 1.88x ahead at 50,000,000 rows in this run, but 0.36x to 0.97x at 250,000 and 500,000 rows in the sweep; the fit wants a bucket ahead by 15% from its crossover up, with 1.5x of headroom, and no such row exists for these.
Sort with a String column, by an int64 keyBehind: 0.89x at 50,000,000 rows, 244.77 ms against Polars 2.0.0's 217.88. To improve.
Filter, then sort by (int32 asc, nullable Float64 desc)Ahead, 2.75x, 317.71 ms against 874.34. Left anyway: it shares a class with the String-column sort, and a class is taken only when every case measuring it is ahead.
One query newly takenGroup-by count over one key at 1,000 groups, 1.91x at 50,000,000 rows, the 2.x table's one addition.
The 63 queries the 1.x table took on 2.0.0 (polars_engine_bench_2026-10-08-polars2-after.csv) against the 52 the 2.x table takes (the refit run), both 50,000,000 rows, best of 5.

The ten group-bys are the honest part. Nine were ahead in this run, some by 1.8x, and the default still leaves them, because the rule is not "ahead at 50,000,000 rows" but "ahead from a crossover, with headroom", and at 250,000 and 500,000 rows Polars 2.0.0 does a two-key sum over 1,000 groups in a few milliseconds the GPU cannot match. MetalEngine(shapes="all") takes every translatable subtree of a million rows or more; either way the report names the rule behind each node. From my machine on 8 October, Polars 2.0.0, a 50,000,000-row frame with a permuted int64 key:

This is what the engine printed on my machine on 8 October, Polars 2.0.0, a 50,000,000-row frame with a permuted int64 key:

import polars as pl, arrowmetal as am          # polars 2.0.0
engine = am.MetalEngine()                      # the measured defaults, the 2.x table
df = lf.sort("id").collect(engine=engine)      # the frame lf.collect() returns
print(engine.last_report)

MetalEngine report for collect (polars 2.0.0, IR (15, 2))
  metal:  Sort#1 [Sort > DataFrameScan] over 50,000,000 rows, ran in 73.09 ms -> 50,000,000 rows
          rule: 50,000,000 input rows is at or above the 10,000,000-row crossover for sort
                (engine table, Benchmarks/results/polars_engine_crossover_2026-10-08-polars2.csv)

# the same plan: Polars 2.0.0 collect() 199.74 ms (streaming), engine="in-memory" 410.51 ms;
# df.equals(lf.collect()) -> True

And the two kinds of no:

lf.group_by("region").agg(pl.col("amount").mean()).collect(engine=engine)

MetalEngine report for collect (polars 2.0.0, IR (15, 2))
  nothing ran on Metal
  polars: GroupBy#1: rule: estimated 50 groups over (region), a Chao1 estimate from a
          512-row sample, 50 to 50: below the measured band for group_by:mean at
          50,000,000 input rows (taken at 3,163 to 3,162,277 groups; ...)
  groups: GroupBy#1 over (region): 50 groups ... (probed in 94 us)

engine.profile(lf)
NotImplementedError: MetalEngine.profile needs LazyFrame.profile, which polars 2.0.0
does not have (Polars 2.0 removed it).

The answers, to the bit

Every one of the 52 taken queries in the refit run is marked equal to Polars'. Counts, keys, join rows and sort orders match exactly; Polars sorts floats in a total order with its null placement per key, and the GPU sort takes both as options on its one radix pass. Float64 group sums are the one place a GPU and a CPU can legitimately differ, in the last bit, because they add in a different order. On a two-key sum of 50,000,000 uniform Float64 values into 100,000 groups the two agree exactly on 11,917 groups and within 2.2e-15 relative on the rest; on the one group I checked against an exactly rounded reference, the GPU's sum is the correctly rounded one.

Install, honestly

  • pip install "arrowmetal[polars]" today installs 0.4.0, which pins Polars to 1.44. The 2.0.0 support and the second table are on main as of 8 October, in the CHANGELOG under "Unreleased", and go out in the next release; until then it is a source build, three commands in docs/POLARS.md, and the extra becomes polars>=1.44,<2.1.
  • MetalEngine.profile raises on 2.0.0; the rest of the API is the same on both majors, and python -m arrowmetal.polars_engine check names the Polars, the IR and the table in force.
  • macOS 14 or later on Apple silicon; every number here is one M4 Max on AC power, load average logged. python -m arrowmetal.bench measures yours in under 30 s.

ArrowMetal is an independent Apache-2.0 project that implements Apache Arrow; Apache Arrow is a trademark of the Apache Software Foundation, and Polars is its maintainers' project, of which I am a user. The table gets refitted again when 2.1 moves the numbers.

FAQ: Polars 2.0, the GPU and Apple silicon

Does Polars 2.0 run on the Apple silicon GPU?

With ArrowMetal's MetalEngine, the subtrees it was measured ahead on do: full sorts, unique, joins and group-bys judged by their group count; Polars runs the rest. The 2.0.0 support is on main, in the next release.

Is Polars 2.0 faster than Polars 1.44?

On an Apple M4 Max at 50,000,000 rows, yes: over 62 group-by, join, sort and unique queries its time is 0.72x of 1.44.1's at the median, 0.25x on unique over two keys. collect() now runs the streaming engine.

How much faster is the GPU than Polars 2.0?

The default takes 52 of 107 benchmark queries at 50,000,000 rows and is 1.35x to 7.94x faster than the faster Polars 2.0.0 engine in every one, none behind.

Why does the GPU engine hand 12 queries back on Polars 2.0?

Polars 2.0.0 got faster on them, so the crossover table has no row where the GPU is ahead by its 15% margin with headroom. The default is fitted per Polars major and takes only what it measured ahead.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts

DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster
Data Engineering10 min read

DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster

Apache DataFusion runs any physical optimizer rule you register. ArrowMetal 0.4.0 ships one for the Apple silicon GPU: the SQL is unchanged, the answers are DataFusion's, and full sorts of 250,000 to 50 million rows run 6.9x to 28.8x faster on an M4 Max. What it takes, what it leaves, and why.

65 views
Read
sqljev: TypeSafe Jev's jev() for SQL Server, Postgres, Snowflake, BigQuery and DuckDB, on Jev or Open-Weight Laya
Data Engineering10 min read

sqljev: TypeSafe Jev's jev() for SQL Server, Postgres, Snowflake, BigQuery and DuckDB, on Jev or Open-Weight Laya

SQL cannot say 'the customer threatens to cancel'. sqljev adds jev(), jev_prob() and jev_choice() to SQL Server, PostgreSQL, MySQL, Snowflake, Databricks, BigQuery, Redshift and DuckDB, and answers them with Laya, an open-weight decision model that runs on your hardware, returns calibrated probabilities instead of text, and can be fine-tuned on your own tables. 140,000 decisions in 271 seconds on one laptop GPU; re-running all 13 queries, 0.9 seconds. Apache 2.0, version 0.1.0.

143 views
Read
pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Data Engineering10 min read

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.

155 views
Read