A 14M-Parameter LLM Keeps a Virtual Fish Tank Alive on an $8 Chip

A distilled LLM runs a fish tank on an $8 ESP32-S3, four Raspberry Pi 5s share a 30B model, and a 2B decision model lands for edge agents.
A Fish Tank's Brain Fits in 7.56 MB
Developer mediacutlet published Pocket Tank on GitHub, a virtual aquarium whose fish are driven by a 14.3-million-parameter language model running entirely on an $8 ESP32-S3. The repo says the model was distilled from a 26-billion-parameter teacher, then quantized and shrunk until it fit a 7.56 MB flash file — small enough that the board's own storage holds both the weights and the simulation around them.
Each fish feeds its tankmates' positions and its own hunger and energy into the model, which picks a behavior — eat, explore, rest, or follow another fish. The ESP32-S3's two cores split the job, one handling inference while the other runs the tank's physics and display, according to CircuitDigest's write-up.
It's a toy, not a product — the fish just need to look alive, not be correct. But a sub-8MB model making real-time behavioral choices on a $8 chip, with no cloud round-trip, is exactly the scale TinyML keeps proving it can handle.
Four Raspberry Pi 5s Share a 30B Model
Hobbyist Hellomatik has published a study showing a cluster of four Raspberry Pi 5 boards running Qwen3-30B-A3B, a 30-billion-parameter mixture-of-experts model, at 15 tokens per second using only the Pis' CPUs, according to LinuxGizmos. The project modifies the existing distributed-llama framework to spread the model's active experts across the boards' network link instead of loading the full weight set onto any single machine.
Fifteen tokens per second is a self-reported figure from one setup, and the mixture-of-experts design means only about 3 billion parameters actually fire per token — the detail that makes the Pi-sized memory footprint possible at all. It's not fast chat, but it's a workable document-reading pace on four boards that together cost less than a mid-range GPU.
Distributed inference across cheap SBCs keeps getting more credible: a model nobody expected outside a datacenter, running on hardware a hobbyist can buy at a maker shop.
Strands Ships a 2B Model Built Only to Decide
Strands Agents released Decider, an open-weight 2-billion-parameter model trained to pick among a fixed set of actions rather than generate free text, on October 7, 2026. The post frames it as a fast, deterministic layer for agent pipelines that hands off to a bigger model only when real reasoning is required. The release reached Hacker News the same day with 280 points and 78 comments.
Strands' own benchmarks, labeled self-reported, claim Decider matches much larger general models on routing and tool-selection tasks. At 2 billion parameters it fits a phone or a small NPU board, though it's still above true microcontroller territory — one more entry in a widening field of narrow decision models that also includes TypeSafe's Jev and llama.cpp's decision-model endpoint.
Whether this becomes a standard layer in edge agent stacks or stays one entrant among several is still open. But the appetite for models that choose instead of chat is specifically an edge-hardware story: small, fast, and cheap enough to run locally.
IBASE Packs Up to 180 TOPS Into a 3.5-Inch Board
IBASE announced the IB966 on October 7, 2026, a 3.5-inch single-board computer built around Intel's Core Ultra Series 3 "Panther Lake" processors, with the vendor citing up to 180 TOPS of aggregate AI performance across the CPU, GPU and NPU, according to LinuxGizmos.
That 180 TOPS is IBASE's own combined figure, not a single-engine NPU number, so treat it as a ceiling rather than a real-workload benchmark. Still, Panther Lake landing in a board this small — aimed at embedded vision and industrial AI PCs rather than laptops — shows how fast edge-AI-PC silicon trickles down to board vendors that usually ship years behind the consumer cycle.
A Palm-Sized Quadruped Crowdfunds on ESP32-C3
ZeroWire Robotics opened crowdfunding on October 8, 2026, for Q8botOne, a palm-sized open-source quadruped robot built around an ESP32-C3 microcontroller and ROBOTIS DYNAMIXEL smart actuators, according to CNX Software.
The ESP32-C3 here handles motion control and communication, not inference — there's no vision-language-action policy on board, and the campaign doesn't claim one. That's the honest caveat: this is a well-documented, hackable walking-robot kit, the kind of open hardware that tends to become a testbed for small on-device policies once someone bolts a camera and a model onto it.
From a $8 chip running a fish tank's brain to four Pis sharing a 30-billion-parameter model, the gap between "toy" and "useful" keeps shrinking in both directions.
References & Citations
- mediacutlet — pocket-tank GitHub repo, 2026-10-06 — https://github.com/mediacutlet/pocket-tank
- CircuitDigest — Pocket Tank write-up, 2026-10-08 — https://circuitdigest.com/news/open-source-pocket-tank-runs-a-14-3m-parameter-ai-model-on-an-esp32-s3-to-give-virtual-fish-their-own-brains
- LinuxGizmos — Raspberry Pi 5 cluster + Qwen3-30B-A3B, 2026-10-07 — https://linuxgizmos.com/raspberry-pi-5-cluster-runs-qwen3-30b-a3b-at-15-tokens-s-on-cpu/
- Strands Agents — Decider announcement, 2026-10-07 — https://strandsagents.com/blog/introducing-strands-decider/
- Hacker News — Strands Decider thread, 2026-10-07 — https://news.ycombinator.com/item?id=49987076
- LinuxGizmos — IBASE IB966 announcement, 2026-10-07 — https://linuxgizmos.com/ibase-ib966-3-5-inch-sbc-with-panther-lake-and-up-to-180-tops/
- CNX Software — Q8botOne crowdfunding, 2026-10-08 — https://www.cnx-software.com/2026/10/08/q8botone-a-palm-sized-open-source-quadruped-robot-with-esp32-c3-dynamixel-smart-actuators/
Subscribe to new posts from theaivibe.org
Related Posts

ESP32 Special: Cloud Voices and an E-Ink Reader, No New On-Device LLM Yet
Two ESP32 voice projects lean on cloud AI while a maker's e-ink reader builds the hardware the next on-device model will need.

Meta Opens Muse to ESP32 Makers as a Tiny LLM Learns to Remember a Hidden Object
Meta ships ESP32 and Linux SDKs for its Muse AI, Qualcomm's next phones get 2nm NPUs, and a 68KB brain gives a cheap robot arm memory.

ExecuTorch Hits 1.0 as llama.cpp Learns to Make Instant AI Decisions
PyTorch declares its on-device AI runtime production-ready while llama.cpp adds a decision-model endpoint and two new NPU modules land for edge boards.