A Classifier That Only Picks, a Model That Skips Tokens, and an AI Firm Bets 2.6B Parameters Can Do Agent Work

Liquid AI ships a 2.6B agentic model for edge hardware, a 144M classifier skips text generation entirely, and a new paper shows byte-level LLMs can beat tokenized ones with sc
Liquid AI's LFM2.5 Bets Agentic Work Fits in 2.6B Parameters
Liquid AI released LFM2.5-2.6B on September 26, 2026, a 2.69-billion-parameter model built for tool-use and instruction-following on constrained hardware — the company's pitch runs from Raspberry Pi boards to laptops and robot controllers. Weights ship day-one in GGUF, MLX, and ONNX formats on Hugging Face, so the same checkpoint runs through llama.cpp, Apple's MLX, or an ONNX runtime without conversion.
Liquid AI says LFM2.5 matches instruction-following and agentic tool-use scores against models roughly four times its size — a comparison the company reports itself. No outside lab has reproduced the numbers yet, per coverage of the release.
The timing is pointed: in a week when several labs pushed larger mixture-of-experts models into the open, Liquid AI is arguing the interesting frontier for agent workloads is memory-constrained hardware, not server racks.
A 144M-Parameter Model That Only Picks, Never Writes
Supersonic Labs released Julia 1 on September 26, 2026: a 144.3-million-parameter open-weight model that skips text generation entirely. Feed it a question and a list of options, and it returns a probability score for each — a classifier, not a chatbot — small enough to run on a laptop CPU with no GPU. Weights sit on Hugging Face under Apache 2.0, flagged by the r/LocalLLaMA thread that surfaced it.
The honest caveat: it cannot draft an email or explain a concept — it ranks a fixed candidate list, the job of a classic classifier head wearing a transformer's clothes. Supersonic Labs has announced a hosted API that isn't public yet.
For edge builders, that narrowness is the point. Intent routing and tool selection are common bottlenecks in agent pipelines, and a 144M CPU-only model that does one job well can slot in without competing for a device's limited RAM.
Byte-Level Models Start Slow, Then Beat Their Tokenizer
A team from FAIR — Kalyani Marathe, Artidoro Pagnoni, and six co-authors — posted "Breaking the Token Ceiling" to arXiv on September 11, 2026, the first large-scale comparison of distilled token models against distilled byte models at matched compute, sweeping roughly 1-billion-parameter transformers up to 1 trillion bytes of training data.
Token models win in the low-compute regime but plateau; byte models start worse and eventually overtake them, with the paper's scaling-law extrapolation projecting distilled byte models to beat Llama 3.2-1B by up to 6.5% and Gemma 3-1B by up to 8.1% on downstream tasks, while needing only one-sixth of the training data. Those figures come from fitted curves, not measured device benchmarks, worth flagging before anyone rewrites a tokenizer pipeline.
The edge angle: byte models operate over a 256-symbol vocabulary instead of roughly 100,000 tokens, cutting logit-storage costs to about a fifth and removing the need for top-k truncation — smaller bookkeeping that matters when every megabyte of a checkpoint has to earn its place on a phone or microcontroller.
Espressif's ESP-Mosaico Wraps a RISC-V MCU in a Square AMOLED
LinuxGizmos reported on September 27, 2026 that Espressif's ESP-Mosaico development kit centers on the dual-core RISC-V ESP32-S31 microcontroller behind a 2.16-inch square AMOLED touchscreen, built as an expandable "smart-interaction" platform rather than a one-off demo board.
Details beyond the chip and display are still filtering out — LinuxGizmos' write-up is the first concrete look at the board, and pricing hasn't surfaced yet.
Square AMOLED faces plus RISC-V cores are becoming the default shape for ESP32-class wearables and smart-home panels; a dev kit built around both from Espressif itself signals where the company expects hobbyists to build next.
LILYGO's T-Dongle-C5 Squeezes ESP32-C5 Into a USB Stick
CNX Software reported on September 27, 2026 that LILYGO's new T-Dongle-C5 packs an ESP32-C5 — dual-band Wi-Fi 6, Bluetooth LE, and 802.15.4 radios — into a USB dongle with a microSD card slot and an optional 0.96-inch OLED display, following on from the earlier ESP32-S3-based T-Dongle-S3.
CNX Software's coverage doesn't list a retail price yet, and the board targets network sniffing, home-automation bridges, and small sensor projects rather than anything running a model on-chip.
It's a reminder that most of the edge stack is still plumbing: cheap radios with enough storage to log or relay data are what feed the tiny models covered elsewhere on this beat.
Nothing here runs a chatbot on a coin cell today, but the pieces — a narrower classifier, a tokenizer-free scaling result, and two new radios — are the kind of parts that end up inside next month's on-device demo.
References & Citations
- Liquid AI — LFM2.5-2.6B model card, Hugging Face, Sept 26, 2026 — https://huggingface.co/LiquidAI/LFM2.5-2.6B
- ai-newspaper.com — Liquid AI Releases LFM2.5 2.6B Model for Edge Devices — https://ai-newspaper.com/articles/liquid-ai-launches-lfm2-5/
- Supersonic Labs — Julia-1 model card, Hugging Face — https://huggingface.co/SupersonicLabs/Julia-1
- r/LocalLLaMA — SupersonicLabs/Julia-1 thread, Sept 27, 2026 — https://www.reddit.com/r/LocalLLaMA/comments/1wr8d4d/supersoniclabsjulia1_hugging_face/
- Marathe, Pagnoni, et al. — Breaking the Token Ceiling, arXiv, Sept 11, 2026 — https://arxiv.org/abs/2609.12303
- LinuxGizmos — ESP-Mosaico, Sept 27, 2026 — https://linuxgizmos.com/esp-mosaico-places-dual-core-risc-v-esp32-s31-behind-a-2-16-inch-square-amoled/
- CNX Software — LILYGO T-Dongle-C5, Sept 27, 2026 — https://www.cnx-software.com/2026/09/27/lilygo-t-dongle-c5-an-esp32-c5-usb-dongle-with-microsd-card-slot-optional-0-96-inch-oled/
Subscribe to new posts from theaivibe.org
Related Posts

Wearable AI Chips Land as an ESP32 Board Learns to Run a Full Offline Voice Loop
From a Qualcomm earbud chip to an ESP32-S3 that hears, thinks and speaks with no cloud, edge AI keeps shrinking into pockets and ears.

A $300 GPU Streams a 177B AI Model From an SSD While llama.cpp Learns to Skip Ahead
Community builders push token throughput further this week — via SSD-streamed MoE experts, prompt-lookup drafting, and a wrapper for Apple's built-in on-device LLM.

Needle Threads a Raspberry Pi, an NPU Learns to Move a Robot Arm Fast, and a Biped Joins the LLM Toolkit
A tool-calling model flips switches on a Pi 5, a Qualcomm NPU speeds up a robot arm sevenfold, and a $2,500 biped joins LeRobot.