
A 312K-Parameter LLM Learns to Flip Switches as Pi Prices Climb Again
A tiny GPIO-control model and a Japanese TTS join the ESP32 pile-up while Raspberry Pi raises prices and a Jetson robot chases bubbles.
Tag
47 posts

A tiny GPIO-control model and a Japanese TTS join the ESP32 pile-up while Raspberry Pi raises prices and a Jetson robot chases bubbles.

A full day inside the ESP32 world: chatty microcontroller LLMs, a $5-chip speech model, and two new boards from Espressif's own community.

From a Qualcomm earbud chip to an ESP32-S3 that hears, thinks and speaks with no cloud, edge AI keeps shrinking into pockets and ears.

Community builders push token throughput further this week — via SSD-streamed MoE experts, prompt-lookup drafting, and a wrapper for Apple's built-in on-device LLM.

Liquid AI ships a 2.6B agentic model for edge hardware, a 144M classifier skips text generation entirely, and a new paper shows byte-level LLMs can beat tokenized ones with sc

A tool-calling model flips switches on a Pi 5, a Qualcomm NPU speeds up a robot arm sevenfold, and a $2,500 biped joins LeRobot.

AMD open-sources fused FlashAttention kernels for XDNA NPUs, Qualcomm ships a Linux preview for Snapdragon X2, and fresh research pushes tiny and trillion-scale models toward

An RP2350 draws faces from noise, an ESP32-S3 writes stories, and AI smart glasses squeeze in a 1-bit model — all without a server.

Three ESP32 builds this week show the chip's range — a reactive desk companion, motion-sensing drumsticks, and a scripting OS — none of them running a language model.

Two new speech stacks and a tool-calling model push voice AI fully on-device, while Nordic and Zephyr quietly harden the hardware underneath.

Small llama.cpp builds keep adding Hexagon DSP ops, NVIDIA's TensorRT-Edge-LLM adds Day-0 model support, and a new leaderboard measures tokens per joule.

PrismML compresses a 27B model to 5.9GB, Intel's BITCOS beats the 1.585-bit ternary limit, and a 44M-parameter model claims exact arithmetic on a laptop CPU.

NVIDIA post-trains Cosmos 3 Edge for on-device manipulation, a PKU lab ships a llama.cpp engine for VLA policies, and Jetson's next Orin Nano gets a ship date.

GMKtec, ASUS and Radxa all shipped NPU hardware this week while OpenVINO and a Qualcomm robotics runtime pushed what those chips can actually run.

NVIDIA posts a 6.4x MLPerf edge win on Jetson AGX Thor, a dense TinyStories model skips the flash trick, and a paper splits VLA robots between cloud and a tiny local model.

A one-chip LLM, a Wi-Fi upgrade to Seeed's tiny displays, and Tuya's push to make ESP32 an AI-agent target, not just a Wi-Fi one.

New wearable and phone releases push transcription, gesture control and silent speech fully on-device, while ESP32 and Jetson tooling keeps pace.

llama.cpp shipped two builds in two days, ExecuTorch hit 1.0 with new NPU backends, and LiteRT tuned fp16 kernels for mobile CPUs.

A ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.

A student-built quadruped, a walking-robot policy on a Rockchip SBC, a tiny STM32N6 vision camera, and a $49 TinyML kit all landed this week.
A phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.

An RP2350 chip runs a diffusion model, an ESP32-S3 speaks Japanese, and researchers tackle mobile power and tiny-drone control at the edge.

A modded Sony PSP runs a tiny LLM, a $3 chip learns to speak, and a 2B open model claims agentic skills for phones.

A GitHub toolkit, a 100MB cloning TTS, a Pi-powered dashcam agent, and a phone-to-watch AI rollout — all inference staying on the device.

A new llama.cpp release, an ExecuTorch-powered 30B agent model, a cheap RK3576 vision board, and a DIY Jetson robot dog mark a busy week for edge toolchains.

Fresh arXiv work rethinks quantization strategy and sustainability, an Apple-adjacent paper shrinks the dictation encoder, and new silicon and Jetson guidance round out the ed

A llama.cpp-style engine ports VLA robot policies to Jetson, NVIDIA doubles entry-level robotics compute, and a $10 chip proves the floor of local AI.

Three chipmakers and a mini-PC builder all attack the same problem this week: getting AI inference closer to memory, not just closer to silicon.

Fresh arXiv work tackles MCU vision drift and phone LLM memory pressure, while a solar bird feeder and a desktop WALL-E show the hobbyist edge staying busy.
DeepSeek V4 Flash gains on-device vision on a Mac, a $1 chip draws pictures, and Qualcomm ships new edge silicon ahead of IFA.

An iFLYTEK spin-off open-sources a 1.7B model claiming native million-token context on-device, while a 14MB tool-caller and an open voice-agent LLM push the small-model race f

llama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.

A 1-bit Qwen derivative fits an iPhone, a healing trick beats its own teacher at 4-bit, and a dense LLM limps along on an $8 chip.

A Rockchip-powered duck robot, two offline Raspberry Pi builds, and audits exposing quantization's blind spots and mislabeled GGUF files.

Intel details a 17 TOPS NPU chip built for chiplets, and an industry trend piece argues edge inference is leaving the pilot stage — both light on independent proof so far.

IBM's Granite 4.2 targets edge devices with a 3B open model, while a new RISC-V AI pocket computer ships locked to its own OS fork.

NVIDIA doubles its entry robotics brain, Perplexity moves agents onto local GPUs, and Liquid AI ships a speedup and a benchmark suite for on-device models.

Hearing aids ship dedicated on-device AI chips, new AI glasses land, and real-phone benchmarks show why raw specs don't tell the whole story.

Meta ships an on-device agentic model, an MoE engine claims 753B on one GPU, and researchers find 10 CVEs in a local inference engine.

Liquid AI distills instead of just rounding for 4-bit LFM2.5 checkpoints, as fresh research flags what low-bit quantization costs in memory and multilingual accuracy.

A Jetson Thor robot policy, a GGUF port for VLA models, a 2.6B tool-calling LLM, and a new AMD robotics module — all inference, no data center.

AMD bumps its NPU mini PC to 192GB of unified memory, Qualcomm ships five agentic apps for Snapdragon X, and Korea rethinks its NPU strategy.

FastFlowLM's first stable release puts a robotics policy on Ryzen AI's NPU, while Korea ships a boxed NPU appliance and a Raspberry Pi learns to narrate what it sees.

A Korean telecom sells an all-in-one on-prem LLM box built on a domestic NPU, while ESP32 tinkerers keep shrinking what a model needs to run.

AMD folds a hobbyist NPU runtime into ROCm, Google shows Gemma driving a robot from a Raspberry Pi 5, and a 45M-parameter model books tool calls on a phone.

A Raspberry Pi 5 runs Gemma and vision models split across CPU and GPU, and a 45M-parameter model fits tool-calling into 28MB of RAM.

A new 'Barista' model and a closer look at Google's Per-Layer Embeddings trick show what running an LLM on a microcontroller can and can't do.