Apache Arrow and Apache DataFusion Now List ArrowMetal: Arrow Compute and a DataFusion Optimizer Rule on the Apple Silicon GPU
ArrowMetal is now on Apache Arrow's Powered By page and on Apache DataFusion's integrations list. What the two listings mean, and what the listed thing does as of 0.5.0, with one real run on each side.
ArrowMetal is now on two Apache project pages, and I am going to enjoy this one. Since 8 September it has been on Apache Arrow's Powered By page, alongside the libraries I learned Arrow from. Since today, 9 October, it is on Apache DataFusion's integrations list in the user guide, as datafusion-arrowmetal. Both went in the way every entry on those lists does: a pull request, read by the project's committers, merged. Here is what the two listings mean, and what the listed thing does today, because 0.5.0 shipped yesterday and both entries describe it.
Apple M4 Max, macOS 27; arrowmetal 0.5.0 from PyPI with polars 2.0.0, best of 5, 9 October 2026; the grid from docs/DATAFUSION.md, DataFusion 55.1.0, 1 October 2026.
ArrowMetal is my Apache-2.0 Arrow compute library for the Apple silicon GPU: Arrow buffers stay in unified memory, the GPU reads them in place, and the results come back as Arrow. It has a Polars engine, a DataFusion optimizer rule, DuckDB extensions and seven language bindings. The launch post covers the design; this one is about the listings.
Two listings, and what they mean
Arrow's Powered By page is the Arrow website's list of projects and products using Apache Arrow; DataFusion's is the section of its user guide for community projects that extend DataFusion. Getting on either means a committer read a description of your work, judged it accurate and on topic, and put their name on the merge. For Arrow that took one morning; for DataFusion, six days, one committer's careful review and a maintainer's merge, with a thank-you at the end. Those are the two communities ArrowMetal is built on, and being read and accepted by them is the part I will remember.
| the two listings | Apache Arrow's Powered By page and Apache DataFusion's integrations list |
|---|---|
| What the page is | Arrow: Powered By, the Arrow website's list of projects and products using Apache Arrow, on the page that also carries the trademark guidance. DataFusion: the “Integrations and Extensions” section of the user guide, “community projects that extend DataFusion or provide integrations with other systems”. |
| How an entry gets in | A pull request, read by the project's committers. Arrow: arrow-site #815, opened and merged on the morning of 8 September after one committer's approval. DataFusion: datafusion #26004, opened 3 October, one committer's review, merged by a maintainer on 9 October. |
| What the entry says | Arrow: “Apache Arrow compute on Apple silicon GPUs through Metal. Arrow buffers live in unified memory the GPU reads in place, so the kernels … run without a copy in either direction …” and the seven bindings. DataFusion: “A physical optimizer rule that runs sorts and measured group-bys on the Apple silicon GPU through Metal.” |
| What it means | The project's committers read a description of the thing, judged it accurate and on topic, and merged it under their name. For a one-person project built on their work, that is the review that counts. |
| What it does not mean | An endorsement, a code review, or any status inside the Foundation. ArrowMetal is an independent Apache-2.0 project, happily so; “Apache Arrow” and “Apache DataFusion” are trademarks of the Apache Software Foundation, which is why the name has no “Apache” in it. |
It matters in a practical way too: these are the pages people read when they ask “what runs on Arrow?” or “what extends DataFusion?”, and an accurate line there reaches the right readers. Both descriptions name what the software does and nothing it does not; that is the bar this post holds to. In fairness to both projects, a listing is not an endorsement or a code review, and ArrowMetal remains independent; the table has the texts and dates.
The Arrow side: what the entry describes
The Powered By entry says the kernels run on Arrow buffers in unified memory without a copy either way, and that data crosses through Arrow's C Data, C Stream and C Device Data interfaces. That is the architecture: a pyarrow array, a Polars Series or a DataFusion RecordBatch is handed over as Arrow, the GPU works on the same bytes, and Arrow comes back. Here is the README's first thing to run, from the 0.5.0 wheel on PyPI, then a 50,000,000-row sort through the Polars engine on Polars 2.0.0, with the engine's report.
$ pip install "arrowmetal[polars]" # 0.5.0 from PyPI; polars 2.0.0 comes with the extra
import pyarrow as pa, arrowmetal as am # the README's first thing to run, as installed
col = am.array(pa.array([1, None, 3, 40])) # a pyarrow array crosses in through the C Data interface
big = col.filter_where(">", 2) # GPU
print(big.sum(), big.to_arrow()) # 43 [3, 40]
keys = am.array(pa.array([0, 1, 0, 2], pa.int32()))
print(keys.group_by(3).sum(col).to_arrow()) # [4, null, 40]
import polars as pl # 50,000,000 permuted int64 keys, a Float64 payload
engine = am.MetalEngine() # the measured defaults; on 2.0.0, the 2.x table
print(engine.explain(df.lazy().sort("k")))
MetalEngine report for explain (polars 2.0.0, IR (15, 2))
metal: Sort#1 [Sort > DataFrameScan] over 50,000,000 rows
rule: 50,000,000 input rows is at or above the 10,000,000-row crossover for sort
(engine table, Benchmarks/results/polars_engine_crossover_2026-10-08-polars2.csv)
# best of 5: polars 2.0.0 collect() 242.37 ms; collect(engine=engine) 43.16 ms; frames equal
am.array(pa.array(...)) is the C Data interface at work; nothing is converted. The report names the rule, the row count and the crossover file, so the decision is checkable. The 43 against 242 milliseconds is one M4 Max sorting 50,000,000 permuted int64 keys with a Float64 payload, best of 5, against Polars 2.0.0's own collect(), frames equal. The engine takes the sort from 10,000,000 rows up, where the crossover table puts it; below that, Polars sorts.
The DataFusion side: what the entry describes
The integrations entry calls the crate a physical optimizer rule that runs sorts and measured group-bys on the Apple silicon GPU through Metal. Such a rule is a hook DataFusion gives to anyone: it sees the physical plan and may replace nodes. This one replaces a full ORDER BY, and the aggregate shapes its measured table takes, with a node that runs the plan on the GPU and hands back the RecordBatches DataFusion expected. The SQL and every other node are untouched. The crate's quickstart, built from crates.io 0.5.0 against DataFusion 55.1.0 on this M4 Max: the sort it takes, the two it leaves.
# Cargo.toml: datafusion = "=55.1.0", datafusion-arrowmetal = "0.5.0"
let rule = ArrowMetalRule::new(ArrowMetalConfig::default()); // full sorts from 250,000 rows
let ctx = session_context(SessionConfig::new().with_target_partitions(4), rule.clone());
ctx.register_table("sales", mem_table)?; // 1,000,000 rows: region, amount, name
== SELECT name, region, amount FROM sales ORDER BY amount DESC NULLS LAST, name
ProjectionExec: expr=[name@2 as name, region@0 as region, amount@1 as amount]
MetalExec: sort=[amount DESC NULLS LAST, name ASC NULLS LAST]
DataSourceExec: partitions=1, partition_sizes=[1]
TAKEN SortExec -- input rows 1000000 (exact) vs min_rows 250000
1000000 rows; the first 3: n001395 | 0 | 1429.4285714285713 ...
== SELECT name, amount FROM sales ORDER BY amount DESC NULLS LAST LIMIT 3
LEFT SortExec: TopK(fetch=3) -- top-k (sort with fetch 3) disabled in config
== SELECT region, count(*) AS n, avg(amount) AS mean FROM sales GROUP BY region ORDER BY region
LEFT AggregateExec: gby=[region], aggr=[count(Int64(1)), avg(sales.amount)]
-- count + sum_avg_f64 over 1 i64 key: the measured table takes it at no
group count and size (results/datafusion_groupby_sweep_2026-10-01.csv)
13 rows; the first 3: 0 | 76924 | 714.7095888611585 ...
The numbers behind “measured” are in the DataFusion post and docs/DATAFUSION.md: on the M4 Max, full sorts from 250,000 to 50,000,000 rows ran 6.9x to 28.8x faster than DataFusion alone across four key types, 50,000,000 rows sorted in 126 CPU-milliseconds against 6,104 on 16 partitions, and 13,632 query pairs, rule off and on, matched with 0 mismatches. Measured on 0.4.0 in late September; 0.5.0 changes the aggregate side, below, not the sort numbers.
What 0.5.0 changed since those two posts
Yesterday's release is the one both entries now describe, and it is the best one yet. Four lines from the release notes matter here.
| ArrowMetal 0.5.0, 8 October 2026 | what changed since the two posts the listings point at, from the release notes |
|---|---|
| Polars 2.0.0 | The Polars engine runs on 2.0.0 as on 1.44, with one measured crossover table per Polars major. On 2.0.0 the default takes 52 of 107 benchmarked queries at 50,000,000 rows on the M4 Max, every one 1.35x to 7.94x faster than the faster Polars 2.0.0 engine, none behind; 12 are handed back. The pip extra is now polars>=1.44,<2.1; the re-measurement post has the cases. |
| Integer mean in one pass | A mean over an integer column runs in one pass in the Polars and plan engines, same bits as before: two int32 keys, 200 groups, 20.7 to 14.3 ms at 50,000,000 rows on polars 2.0.0, per the release notes. |
| DataFusion crate | datafusion-arrowmetal 0.5.0 on DataFusion 55.1: full ORDER BY sorts, and by default the aggregate shapes its measured table takes: count(*) and DISTINCT over two int32 keys, integer MIN/MAX over two integer keys or one int64 key, on a MemTable of 50,000,000 rows or more. Hash joins are translated; the measured join table takes none yet. |
| Build targets | DuckDB extensions built and tested against DuckDB 1.5.6; macOS 14 or later on Apple silicon; the Rust crate's build script stops the build on any other target. |
The Polars line closes a caveat from Wednesday's post, which said the 2.0.0 support was on main, unreleased. As of 0.5.0 it is the wheel on PyPI, which is why the card above needed only a fresh venv.
Install
- Python:
pip install "arrowmetal[polars]"installs 0.5.0 and a Polars from 1.44 to 2.0;python -m arrowmetal.polars_engine checkprints the Polars, the IR and the table in force. Without the extra, pyarrow is the only dependency. - Rust:
datafusion = "=55.1.0"anddatafusion-arrowmetal = "0.5.0"inCargo.toml; the crate links the Metal library throughARROWMETAL_LIBor the copy the Python wheel installs. docs.rs and docs/RUST.md have the finding rules. - macOS 14 or later on Apple silicon only. Every number here is one M4 Max on AC power;
python -m arrowmetal.benchmeasures yours in under 30 s.
ArrowMetal is an independent Apache-2.0 project; Apache Arrow and Apache DataFusion are trademarks of the Apache Software Foundation, whose projects I use and build on every day. My thanks to the committers on both sides who read the entries and let them in. Two gates, one small craft, a harbour I am glad to be moored in.
FAQ: the listings, Apache Arrow, Apache DataFusion and the Apple silicon GPU
Is ArrowMetal an Apache project?
No, and proudly independent: an Apache-2.0 project that implements Apache Arrow and extends Apache DataFusion. Both listings are entries on community-maintained lists, read and merged by the projects' committers; they are a welcome, not an endorsement.
Does Apache DataFusion ship the GPU rule?
No. datafusion-arrowmetal is a separate crate on crates.io, pinned to DataFusion 55.1.0. You register its rule on a SessionContext; the SQL and every other plan node stay DataFusion's.
Which Macs does it run on?
Apple silicon on macOS 14 or later. The Python wheel carries the Metal library; the Rust crate finds it through ARROWMETAL_LIB or the wheel's copy. The build script stops on any other target.
Are the answers the same as the CPU's?
The differential grid runs 13,632 DataFusion query pairs, rule off and on, with 0 mismatches; every Polars benchmark query is checked equal to Polars' own frame. Float64 runs through a software IEEE-754 path, bit-exact with the CPU.
Read next
References & Citations
- Apache Arrow (2026). Powered By: the ArrowMetal entry, added by apache/arrow-site #815, merged 8 September 2026.
- Apache DataFusion (2026). User Guide, Introduction: Integrations and Extensions: the datafusion-arrowmetal entry, added by apache/datafusion #26004, merged 9 October 2026.
- ArrowMetal (2026). v0.5.0 release notes, CHANGELOG and README at the tag; docs/DATAFUSION.md for the sort numbers (M4 Max, DataFusion 55.1.0, 29 September 2026) and the 13,632-pair grid (1 October 2026); docs/POLARS.md for the Polars 2.0.0 default.
- arrowmetal 0.5.0 on PyPI (polars extra
>=1.44,<2.1) and datafusion-arrowmetal 0.5.0 on crates.io (DataFusion=55.1.0). - Apache Software Foundation. Trademark policy.
- The two code cards are from runs for this post on 9 October 2026 (arrowmetal 0.5.0 wheel from PyPI in a fresh venv, polars 2.0.0, Python 3.13, Apple M4 Max, macOS 27.0, 9 October 2026; the Rust card from the crate's quickstart example built against crates.io 0.5.0); report lines are the tools' actual output, trimmed for width.
Subscribe to new posts from theaivibe.org
Related Posts
Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still
Polars 2.0.0 shipped on 6 October and runs its group-bys, joins and sorts faster than 1.44. Two days later ArrowMetal's MetalEngine was re-measured against it on an Apple M4 Max: the GPU default takes 52 of 107 queries at 50 million rows, each 1.35x to 7.94x faster than the faster Polars 2.0 engine, and hands 12 back that it used to take. What Polars 2.0 changed, and how a crossover table answers it.
DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster
Apache DataFusion runs any physical optimizer rule you register. ArrowMetal 0.4.0 ships one for the Apple silicon GPU: the SQL is unchanged, the answers are DataFusion's, and full sorts of 250,000 to 50 million rows run 6.9x to 28.8x faster on an M4 Max. What it takes, what it leaves, and why.

sqljev: TypeSafe Jev's jev() for SQL Server, Postgres, Snowflake, BigQuery and DuckDB, on Jev or Open-Weight Laya
SQL cannot say 'the customer threatens to cancel'. sqljev adds jev(), jev_prob() and jev_choice() to SQL Server, PostgreSQL, MySQL, Snowflake, Databricks, BigQuery, Redshift and DuckDB, and answers them with Laya, an open-weight decision model that runs on your hardware, returns calibrated probabilities instead of text, and can be fine-tuned on your own tables. 140,000 decisions in 271 seconds on one laptop GPU; re-running all 13 queries, 0.9 seconds. Apache 2.0, version 0.1.0.