Back to Edge

Seven $5 Chips Learn to Share an LLM, One Bit at a Time

Prateek SinghOctober 4, 20265 min read6 views
Seven $5 Chips Learn to Share an LLM, One Bit at a Time

A BitNet cluster splits an LLM across seven ESP32-S3 boards, a 2-bit tool-calling model lands on GitHub, and a quantization paper gets llama.cpp 15x faster on an M4 Pro.

A BitNet Cluster Splits an LLM Across Seven ESP32-S3 Boards

Developer Low Zi Hong published a project on GitHub on September 26, 2026, that runs a pruned Qwen2-0.5B model across seven ESP32-S3 microcontrollers linked by an SPI daisy-chain. One board handles tokenizing and sampling; the other six each carry four transformer layers using Microsoft's BitNet 1.58-bit ternary scheme, which packs weights of {-1, 0, +1} into two bits apiece. Total draw during inference is roughly 1.5 watts, dropping to about 1.15 watts idle, with each node taking around 1.3 seconds per step, according to the write-up by Shamyl Bin Mansoor.

The honest caveat: this is a proof of concept, not a usable assistant. The model is undertrained and the output is described as only semi-coherent. It leans on Microsoft's BitNet b1.58 2B4T release from April 2025, the first native 1-bit LLM trained from scratch at 2B parameters.

What it demonstrates is still worth noting — ternary weights shrink a 0.5B model to around 100MB, too big for one ESP32-S3's 16MB flash but small enough to distribute across a handful of $5 boards, which is a real architectural path for microcontroller clusters even if nobody is shipping a product on it yet.

A 2-Bit Model Built Just for Tool Calls Lands on GitHub

Cactus Compute pushed Needle, a 2-bit quantized model it calls an "automation foundation model for tiny devices," to the top of GitHub's trending feed on October 4, 2026, with checkpoints sized between 8MB and 29MB. The repo frames it around tool calls and structured data extraction rather than open conversation — the kind of narrow job a microcontroller actually needs, according to the cactus-compute/needle repository.

No independent benchmark numbers accompany the launch yet, so treat the size claims as self-reported until someone runs it on real hardware and times it. The project has pulled in over 13,000 stars since being created in February 2026, which signals interest more than it proves accuracy.

If it holds up, a sub-30MB model that reliably picks the right function and fills the right JSON fields is more useful on a microcontroller than a bigger model that merely talks — tool-calling at this size is exactly the gap most TinyML stacks still have.

A Distillation Recipe Pushes llama.cpp to 15x on an M4 Pro

A new arXiv paper on EdgeRazor describes a mixed-precision, quantization-aware distillation framework that claims its 1.88-bit Qwen3-0.6B variant beats state-of-the-art 2-bit baselines by 11.27 points and 3-bit baselines by 4.38 points across 14 tasks. The team reports that a 1.58-bit version cuts storage from 1.11GB to 0.19GB and, run through llama.cpp on an Apple M4 Pro, decodes 15.16 times faster than the 16-bit baseline, per the EdgeRazor paper.

These are the authors' own numbers, not an independent replication, and the paper's training-budget savings (4-10x lower than leading QAT methods) are self-reported too. The method is weight-and-activation quantization, which is harder to pull off than weight-only schemes.

Still, a sub-200MB model decoding over 15x faster on consumer Apple silicon is the kind of result that matters for phones and laptops doing on-device inference without a discrete GPU — assuming it survives scrutiny once the code lands.

Efinix Brings 64-Bit RISC-V to FPGAs for Embedded Linux and Edge AI

Efinix announced the Sapphire RV64 this week, a configurable 64-bit RISC-V soft SoC meant to run on the company's FPGA fabric for embedded Linux and edge AI workloads that need more memory headroom than a bare microcontroller core, according to LinuxGizmos.

It's a step up from Efinix's existing 32-bit Sapphire soft cores, which were never meant to boot a full Linux stack. The RV64 variant is aimed squarely at that gap — a synthesizable core a designer can drop into an FPGA alongside their own AI accelerator logic rather than buying a fixed SoC.

The appeal for edge AI is flexibility: a soft core that can run Linux and sit next to custom silicon on the same FPGA fabric opens the door to purpose-built inference boards without committing to a specific chip vendor's roadmap.

New Firmware Turns an ESP32 Into a Dual-Band Software-Defined Radio

A project called ESP-SDR, covered by CNX Software on October 4, 2026, repurposes the ESP32's existing radio hardware to receive on both the 2.4GHz and 5GHz bands as a software-defined radio, without any extra RF front-end chip. Hackaday's companion write-up frames it alongside the RTL-SDR lineage of chips that got a second life once someone found an undocumented receive path.

This is firmware, not a model — there's no inference claim here, and the dual-band range depends on what the ESP32's radio silicon can actually tune to cleanly, which the write-ups don't fully pin down yet.

For the maker side of this beat, it's a reminder that the ESP32's value keeps expanding past Wi-Fi and Bluetooth: a $3 chip doing double duty as a spectrum analyzer is the same instinct that put tiny LLMs on these boards in the first place.

Five stories, two bits per weight in three of them — the quantization race keeps finding new places to put a model, from a seven-board microcontroller cluster to an FPGA soft core that just wants to boot Linux.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts