Tag

on-device inference

6 posts

Ternary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AI
Edge AI4 min read

Ternary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AI

A ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.

66 views
Read
Qualcomm Chases 30B AI Models on a Phone as Edge Silicon Keeps Multiplying
Edge AI4 min read

Qualcomm Chases 30B AI Models on a Phone as Edge Silicon Keeps Multiplying

A phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.

116 views
Read
llama.cpp Teaches an NPU to Share the Load as Small AI Engines Keep Multiplying
Edge AI4 min read

llama.cpp Teaches an NPU to Share the Load as Small AI Engines Keep Multiplying

llama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.

104 views
Read
A 27B LLM Shrinks to Phone Size as Edge Quantization Keeps Rewriting the Rules
Edge AI4 min read

A 27B LLM Shrinks to Phone Size as Edge Quantization Keeps Rewriting the Rules

A 1-bit Qwen derivative fits an iPhone, a healing trick beats its own teacher at 4-bit, and a dense LLM limps along on an $8 chip.

72 views
Read
Edge Dispatch: Jetson Orin Nano 2 Doubles Edge Robotics AI, While Local LLM Agents Get a Bigger Home on the Desktop
Edge AI4 min read

Edge Dispatch: Jetson Orin Nano 2 Doubles Edge Robotics AI, While Local LLM Agents Get a Bigger Home on the Desktop

NVIDIA doubles its entry robotics brain, Perplexity moves agents onto local GPUs, and Liquid AI ships a speedup and a benchmark suite for on-device models.

104 views
Read
Edge Dispatch: Korea's KT Ships a Boxed NPU LLM Station While the ESP32 Crowd Trims Memory Further
Edge AI3 min read

Edge Dispatch: Korea's KT Ships a Boxed NPU LLM Station While the ESP32 Crowd Trims Memory Further

A Korean telecom sells an all-in-one on-prem LLM box built on a domestic NPU, while ESP32 tinkerers keep shrinking what a model needs to run.

121 views
Read