Skip to content

Tag

Apache Arrow

11 posts

Apache Arrow and Apache DataFusion Now List ArrowMetal: Arrow Compute and a DataFusion Optimizer Rule on the Apple Silicon GPU
Data Engineering10 min read

Apache Arrow and Apache DataFusion Now List ArrowMetal: Arrow Compute and a DataFusion Optimizer Rule on the Apple Silicon GPU

ArrowMetal is now on Apache Arrow's Powered By page and on Apache DataFusion's integrations list. What the two listings mean, and what the listed thing does as of 0.5.0, with one real run on each side.

49 views
Read
Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still
Data Engineering10 min read

Polars 2.0 Is Faster. With ArrowMetal on the Apple Silicon GPU, 52 of 107 Queries Go 1.35x to 7.94x Faster Still

Polars 2.0.0 shipped on 6 October and runs its group-bys, joins and sorts faster than 1.44. Two days later ArrowMetal's MetalEngine was re-measured against it on an Apple M4 Max: the GPU default takes 52 of 107 queries at 50 million rows, each 1.35x to 7.94x faster than the faster Polars 2.0 engine, and hands 12 back that it used to take. What Polars 2.0 changed, and how a crossover table answers it.

51 views
Read
DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster
Data Engineering10 min read

DataFusion on the Apple Silicon GPU: One Optimizer Rule, the Same SQL, Sorts 6.9x to 28.8x Faster

Apache DataFusion runs any physical optimizer rule you register. ArrowMetal 0.4.0 ships one for the Apple silicon GPU: the SQL is unchanged, the answers are DataFusion's, and full sorts of 250,000 to 50 million rows run 6.9x to 28.8x faster on an M4 Max. What it takes, what it leaves, and why.

98 views
Read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Data Engineering19 min read

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.

160 views
Read
Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway
Data Engineering15 min read

Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway

The Metal Shading Language has no double type, and float64 is the default number in Python, pandas and Apache Arrow. Building ArrowMetal meant getting past three walls: a GPU with no 64-bit floats, a missing 64-bit atomic add, and a Swift compiler bug that reports errors nobody threw. Here is how each one was solved, what it cost, and why a GPU that cannot add two doubles sorts 50,000,000 of them in 32 ms.

235 views
Read
Polars vs DuckDB vs ArrowMetal GPU on Apple Silicon: Sort and Group-By Benchmarks
Data Engineering12 min read

Polars vs DuckDB vs ArrowMetal GPU on Apple Silicon: Sort and Group-By Benchmarks

Polars, DuckDB and ArrowMetal on an Apple M4 Max: sort and group-by benchmarks at 10M and 50M rows, wall time next to CPU time, and the rows where the CPU is still ahead.

255 views
Read
Apache Arrow Compute on the Apple Silicon GPU: The First Arrow Project I Created That Does It, With 173 Operations Measured Against Polars, pyarrow and pandas
Data Engineering11 min read

Apache Arrow Compute on the Apple Silicon GPU: The First Arrow Project I Created That Does It, With 173 Operations Measured Against Polars, pyarrow and pandas

I built ArrowMetal, the first Apache Arrow project I could find that runs compute on the Apple silicon GPU. Apple silicon has one memory for CPU and GPU, and an Arrow buffer in shared Metal memory is already a GPU buffer; no Arrow project used that. ArrowMetal does: 307 of Arrow's 307 compute functions, seven languages, take at 24.2x pyarrow on an M4 Max, and 339 benchmark rows against the fastest CPU idiom of Polars, pyarrow, pandas and numpy, including the 77 where the CPU is still ahead.

293 views
Read
“PostgreSQL-compatible” Is Not PostgreSQL: What Arrow's Native ADBC Driver Does on 14 Wire-Compatible Databases
Data Engineering12 min read

“PostgreSQL-compatible” Is Not PostgreSQL: What Arrow's Native ADBC Driver Does on 14 Wire-Compatible Databases

I ran the native PostgreSQL and MySQL ADBC drivers against 28 databases that speak their protocols. Half stopped. Then I found a bug in my own driver.

153 views
Read
Apache Arrow ADBC Just Listed My ODBC Bridge on Its Official Integrations Page — Seven Days After v0.1.0
Data Engineering6 min read

Apache Arrow ADBC Just Listed My ODBC Bridge on Its Official Integrations Page — Seven Days After v0.1.0

The Apache Arrow ADBC documentation now lists adbcBridge on its Tools & Integrations page — seven days after I released v0.1.0. I filed the listing request on August 29; on August 31 a project member invited a PR, and it was merged six hours after the invitation. What the entry says, how the week that earned it went, and what it changes for anyone with an ODBC-only database.

172 views
Read
Your ODBC Driver Says SQL_SUCCESS and Lies: 24 Bugs Found in 13 Database Projects (adbcBridge Part 2)
Data Engineering15 min read

Your ODBC Driver Says SQL_SUCCESS and Lies: 24 Bugs Found in 13 Database Projects (adbcBridge Part 2)

Part 2 of the adbcBridge story. Running one workload through 46 databases on three operating systems turned up 24 defects that belong to other projects — twelve of them return wrong or lost data under SQL_SUCCESS. Every one is filed upstream with a reproduction that needs no adbcBridge in the stack. Here is the ledger, what the bugs have in common, and what happened when the maintainers read them: within the first week, two fixes landed upstream and four more fix PRs went up.

171 views
Read
I Built an Apache Arrow ADBC Driver for Every ODBC Database: 46 Databases, 5 Languages, Every Number Measured
Data Engineering21 min read

I Built an Apache Arrow ADBC Driver for Every ODBC Database: 46 Databases, 5 Languages, Every Number Measured

Native Apache Arrow ADBC drivers exist for a handful of databases. The other few hundred ship an ODBC driver and nothing else. adbcBridge is one plain-C11 shared library that turns every ODBC driver on your machine into an Arrow-native ADBC driver — columnar record batches out, bulk ingest in — from Python, Rust, Go, Java and C#. Today it is public: 46 databases verified on Linux, 41 on macOS, 45 on Windows, five languages measured against all of them, and every figure named with the laptop and the load it was taken under.

402 views
Read