
Ternary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AI
A ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.
Tag
6 posts

A ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.
A phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.

llama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.

A 1-bit Qwen derivative fits an iPhone, a healing trick beats its own teacher at 4-bit, and a dense LLM limps along on an $8 chip.

NVIDIA doubles its entry robotics brain, Perplexity moves agents onto local GPUs, and Liquid AI ships a speedup and a benchmark suite for on-device models.

A Korean telecom sells an all-in-one on-prem LLM box built on a domestic NPU, while ESP32 tinkerers keep shrinking what a model needs to run.