Tag

quantization

9 posts

Seven $5 Chips Learn to Share an LLM, One Bit at a Time
Edge AI5 min read

Seven $5 Chips Learn to Share an LLM, One Bit at a Time

A BitNet cluster splits an LLM across seven ESP32-S3 boards, a 2-bit tool-calling model lands on GitHub, and a quantization paper gets llama.cpp 15x faster on an M4 Pro.

105 views
Read
A Classifier That Only Picks, a Model That Skips Tokens, and an AI Firm Bets 2.6B Parameters Can Do Agent Work
Edge AI4 min read

A Classifier That Only Picks, a Model That Skips Tokens, and an AI Firm Bets 2.6B Parameters Can Do Agent Work

Liquid AI ships a 2.6B agentic model for edge hardware, a 144M classifier skips text generation entirely, and a new paper shows byte-level LLMs can beat tokenized ones with sc

82 views
Read
NPUs Learn to Fuse: AMD Opens Its XDNA Compiler, Qualcomm Previews Linux, and Two New On-Device AI Models Land
Edge AI5 min read

NPUs Learn to Fuse: AMD Opens Its XDNA Compiler, Qualcomm Previews Linux, and Two New On-Device AI Models Land

AMD open-sources fused FlashAttention kernels for XDNA NPUs, Qualcomm ships a Linux preview for Snapdragon X2, and fresh research pushes tiny and trillion-scale models toward

137 views
Read
A 27B AI Model Shrinks to 5.9GB as Ternary Quantization Keeps Pushing the Floor Down
Edge AI5 min read

A 27B AI Model Shrinks to 5.9GB as Ternary Quantization Keeps Pushing the Floor Down

PrismML compresses a 27B model to 5.9GB, Intel's BITCOS beats the 1.585-bit ternary limit, and a 44M-parameter model claims exact arithmetic on a laptop CPU.

94 views
Read
Ternary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AI
Edge AI4 min read

Ternary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AI

A ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.

95 views
Read
Quantization Papers Pile Up as Researchers Argue Over Where Small AI Should Spend Its Bits
Edge AI4 min read

Quantization Papers Pile Up as Researchers Argue Over Where Small AI Should Spend Its Bits

Fresh arXiv work rethinks quantization strategy and sustainability, an Apple-adjacent paper shrinks the dictation encoder, and new silicon and Jetson guidance round out the ed

132 views
Read
A 27B LLM Shrinks to Phone Size as Edge Quantization Keeps Rewriting the Rules
Edge AI4 min read

A 27B LLM Shrinks to Phone Size as Edge Quantization Keeps Rewriting the Rules

A 1-bit Qwen derivative fits an iPhone, a healing trick beats its own teacher at 4-bit, and a dense LLM limps along on an $8 chip.

80 views
Read
A Duck Robot Runs Its Balance Loop On-Chip as Local AI Quantization Claims Get Audited
Edge AI4 min read

A Duck Robot Runs Its Balance Loop On-Chip as Local AI Quantization Claims Get Audited

A Rockchip-powered duck robot, two offline Raspberry Pi builds, and audits exposing quantization's blind spots and mislabeled GGUF files.

103 views
Read
Edge Dispatch: Liquid AI Trains Its Way Around the Q4_0 Quality Tax, While New Papers Chip Away at What Quantization Quietly Breaks
Edge AI4 min read

Edge Dispatch: Liquid AI Trains Its Way Around the Q4_0 Quality Tax, While New Papers Chip Away at What Quantization Quietly Breaks

Liquid AI distills instead of just rounding for 4-bit LFM2.5 checkpoints, as fresh research flags what low-bit quantization costs in memory and multilingual accuracy.

123 views
Read