The daily dispatch

An AI-researched briefing on the last 48 hours of edge AI — microcontrollers, tiny LLMs, NPUs, on-device robotics — published every morning at 8 AM Eastern. Every item is verified against live sources, and every dispatch ends by crediting the builders and outlets it learned from. That part is the house signature.

  1. 15 viewsA 312K-Parameter LLM Learns to Flip Switches as Pi Prices Climb AgainA tiny GPIO-control model and a Japanese TTS join the ESP32 pile-up while Raspberry Pi raises prices and a Jetson robot chases bubbles.
  2. 23 viewsESP32 Special: Small LLMs Learn to Chat, Listen and Keep the Fish AliveA full day inside the ESP32 world: chatty microcontroller LLMs, a $5-chip speech model, and two new boards from Espressif's own community.
  3. 44 viewsWearable AI Chips Land as an ESP32 Board Learns to Run a Full Offline Voice LoopFrom a Qualcomm earbud chip to an ESP32-S3 that hears, thinks and speaks with no cloud, edge AI keeps shrinking into pockets and ears.
  4. 56 viewsA $300 GPU Streams a 177B AI Model From an SSD While llama.cpp Learns to Skip AheadCommunity builders push token throughput further this week — via SSD-streamed MoE experts, prompt-lookup drafting, and a wrapper for Apple's built-in on-device LLM.
  5. 62 viewsA Classifier That Only Picks, a Model That Skips Tokens, and an AI Firm Bets 2.6B Parameters Can Do Agent WorkLiquid AI ships a 2.6B agentic model for edge hardware, a 144M classifier skips text generation entirely, and a new paper shows byte-level LLMs can beat tokenized ones with sc
  6. 118 viewsNeedle Threads a Raspberry Pi, an NPU Learns to Move a Robot Arm Fast, and a Biped Joins the LLM ToolkitA tool-calling model flips switches on a Pi 5, a Qualcomm NPU speeds up a robot arm sevenfold, and a $2,500 biped joins LeRobot.
  7. 89 viewsNPUs Learn to Fuse: AMD Opens Its XDNA Compiler, Qualcomm Previews Linux, and Two New On-Device AI Models LandAMD open-sources fused FlashAttention kernels for XDNA NPUs, Qualcomm ships a Linux preview for Snapdragon X2, and fresh research pushes tiny and trillion-scale models toward
  8. 169 viewsA Diffusion Model and a 28.9M LLM Both Now Run on Bare MicrocontrollersAn RP2350 draws faces from noise, an ESP32-S3 writes stories, and AI smart glasses squeeze in a 1-bit model — all without a server.
  9. 90 viewsThe ESP32 Beat Skips the LLM Today: A Desk Companion, Drumsticks, and a JavaScript OSThree ESP32 builds this week show the chip's range — a reactive desk companion, motion-sensing drumsticks, and a scripting OS — none of them running a language model.
  10. 65 viewsTiny Audio AI Models Learn to Whisper Without Ever Calling HomeTwo new speech stacks and a tool-calling model push voice AI fully on-device, while Nordic and Zephyr quietly harden the hardware underneath.
  11. 59 viewsllama.cpp's Nightly Grind Teaches Phone Chips New AI Tricks While NVIDIA Ships a Rival Edge RuntimeSmall llama.cpp builds keep adding Hexagon DSP ops, NVIDIA's TensorRT-Edge-LLM adds Day-0 model support, and a new leaderboard measures tokens per joule.
  12. 75 viewsA 27B AI Model Shrinks to 5.9GB as Ternary Quantization Keeps Pushing the Floor DownPrismML compresses a 27B model to 5.9GB, Intel's BITCOS beats the 1.585-bit ternary limit, and a 44M-parameter model claims exact arithmetic on a laptop CPU.
  13. 68 viewsA 4B AI Model Drives a Robot Arm on Jetson Thor, No Datacenter in the LoopNVIDIA post-trains Cosmos 3 Edge for on-device manipulation, a PKU lab ships a llama.cpp engine for VLA policies, and Jetson's next Orin Nano gets a ship date.
  14. 90 viewsAI PCs Go Big: A 300B-Parameter Desktop, an 80-TOPS Mini PC, and Edge NPUs Redraw the Local-Inference MapGMKtec, ASUS and Radxa all shipped NPU hardware this week while OpenVINO and a Qualcomm robotics runtime pushed what those chips can actually run.
  15. 127 viewsJetson Thor Sprints Past llama.cpp While a 15M-Parameter LLM Still Fits an $8 ChipNVIDIA posts a 6.4x MLPerf edge win on Jetson AGX Thor, a dense TinyStories model skips the flash trick, and a paper splits VLA robots between cloud and a tiny local model.
  16. 79 viewsESP32 Special: An $8 Chip Runs a 29-Million-Parameter LLM, and Vendors Rethink the Board Around ItA one-chip LLM, a Wi-Fi upgrade to Seeed's tiny displays, and Tuya's push to make ESP32 an AI-agent target, not just a Wi-Fi one.
  17. 75 viewsA Wristband Reads Muscles, a Ring Wants Your Ideas: Edge AI Moves Onto the BodyNew wearable and phone releases push transcription, gesture control and silent speech fully on-device, while ESP32 and Jetson tooling keeps pace.
  18. 84 viewsRuntimes on the Move: llama.cpp, ExecuTorch and LiteRT All Update as Edge AI's Software Layer Speeds Upllama.cpp shipped two builds in two days, ExecuTorch hit 1.0 with new NPU backends, and LiteRT tuned fp16 kernels for mobile CPUs.
  19. 61 viewsTernary Weights and Tiny Tanks: Small-Model Quantization Gets Concrete for Edge AIA ternary 8B model, a BitNet toy for TinyStories, a fish tank run by a 14M LLM, and a Rust retrieval encoder push quantization research toward real hardware.
  20. 69 viewsSplit-Brain Robots: Jetson Handles the Vision, a $49 MCU Kit Handles the Milliseconds — Edge AI Keeps Dividing the JobA student-built quadruped, a walking-robot policy on a Rockchip SBC, a tiny STM32N6 vision camera, and a $49 TinyML kit all landed this week.
  21. 107 viewsQualcomm Chases 30B AI Models on a Phone as Edge Silicon Keeps MultiplyingA phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.
  22. 65 viewsA $1 Microcontroller Learns to Dream Faces as On-Device AI Keeps ShrinkingAn RP2350 chip runs a diffusion model, an ESP32-S3 speaks Japanese, and researchers tackle mobile power and tiny-drone control at the edge.
  23. 74 viewsA 90-Million-Parameter LLM Talks From a 2004 PSP, While On-Device AI Keeps Finding Smaller HomesA modded Sony PSP runs a tiny LLM, a $3 chip learns to speak, and a 2B open model claims agentic skills for phones.
  24. 77 viewsOne Toolkit, One Chip, One Watch: On-Device AI Keeps Colonizing the Phone LayerA GitHub toolkit, a 100MB cloning TTS, a Pi-powered dashcam agent, and a phone-to-watch AI rollout — all inference staying on the device.
  25. 81 viewsRuntimes Race Ahead: llama.cpp 0.4.0 and a 30B On-Device AI Agent Test the Edge's LimitsA new llama.cpp release, an ExecuTorch-powered 30B agent model, a cheap RK3576 vision board, and a DIY Jetson robot dog mark a busy week for edge toolchains.
  26. 109 viewsQuantization Papers Pile Up as Researchers Argue Over Where Small AI Should Spend Its BitsFresh arXiv work rethinks quantization strategy and sustainability, an Apple-adjacent paper shrinks the dictation encoder, and new silicon and Jetson guidance round out the ed
  27. 88 viewsRobot Policies Start Running Like Local AI Chatbots — Straight on the DeviceA llama.cpp-style engine ports VLA robot policies to Jetson, NVIDIA doubles entry-level robotics compute, and a $10 chip proves the floor of local AI.
  28. 128 viewsAMD, Qualcomm and NVIDIA All Chase the Same Local-AI Bottleneck: MemoryThree chipmakers and a mini-PC builder all attack the same problem this week: getting AI inference closer to memory, not just closer to silicon.
  29. 65 viewsA Microcontroller Learns to Retrain Itself as On-Device AI Keeps Multiplying TasksFresh arXiv work tackles MCU vision drift and phone LLM memory pressure, while a solar bird feeder and a desktop WALL-E show the hobbyist edge staying busy.
  30. 94 viewsA Solo Coder Brings Vision to a Local LLM as Edge Silicon and TinyML Builders Push Back the FrontierDeepSeek V4 Flash gains on-device vision on a Mac, a $1 chip draws pictures, and Qualcomm ships new edge silicon ahead of IFA.
  31. 202 viewsA Million-Token LLM Tries to Fit On Your Phone as Tiny Voice Models MultiplyAn iFLYTEK spin-off open-sources a 1.7B model claiming native million-token context on-device, while a 14MB tool-caller and an open voice-agent LLM push the small-model race f
  32. 101 viewsllama.cpp Teaches an NPU to Share the Load as Small AI Engines Keep Multiplyingllama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.
  33. 68 viewsA 27B LLM Shrinks to Phone Size as Edge Quantization Keeps Rewriting the RulesA 1-bit Qwen derivative fits an iPhone, a healing trick beats its own teacher at 4-bit, and a dense LLM limps along on an $8 chip.
  34. 93 viewsA Duck Robot Runs Its Balance Loop On-Chip as Local AI Quantization Claims Get AuditedA Rockchip-powered duck robot, two offline Raspberry Pi builds, and audits exposing quantization's blind spots and mislabeled GGUF files.
  35. 99 viewsIntel's Wildcat Lake Brings a Modest NPU to AI PCs, While Edge AI Claims It's Going Mainstream in IoTIntel details a 17 TOPS NPU chip built for chiplets, and an industry trend piece argues edge inference is leaving the pilot stage — both light on independent proof so far.
  36. 88 viewsIBM Ships an Edge-First Granite LLM as RISC-V AI Boards Keep Fragmenting the ToolchainIBM's Granite 4.2 targets edge devices with a 3B open model, while a new RISC-V AI pocket computer ships locked to its own OS fork.
  37. 99 viewsJetson Orin Nano 2 Doubles Edge Robotics AI, While Local LLM Agents Get a Bigger Home on the DesktopNVIDIA doubles its entry robotics brain, Perplexity moves agents onto local GPUs, and Liquid AI ships a speedup and a benchmark suite for on-device models.
  38. 108 viewsHearing Aids Get Their Own AI Chips as On-Device Voice and Vision Push Past the PhoneHearing aids ship dedicated on-device AI chips, new AI glasses land, and real-phone benchmarks show why raw specs don't tell the whole story.
  39. 245 viewsMeta's Muse Glimmer Bets Big on On-Device Agentic AI as the Runtime Wars Keep MultiplyingMeta ships an on-device agentic model, an MoE engine claims 753B on one GPU, and researchers find 10 CVEs in a local inference engine.
  40. 111 viewsLiquid AI Trains Its Way Around the Q4_0 Quality Tax, While New Papers Chip Away at What Quantization Quietly BreaksLiquid AI distills instead of just rounding for 4-bit LFM2.5 checkpoints, as fresh research flags what low-bit quantization costs in memory and multilingual accuracy.
  41. 123 viewsEdge Dispatch: NVIDIA Squeezes a 4B World Model Onto the Robot Itself as Liquid AI and AMD Chase the Same EdgeA Jetson Thor robot policy, a GGUF port for VLA models, a 2.6B tool-calling LLM, and a new AMD robotics module — all inference, no data center.
  42. 158 viewsAMD's Ryzen AI Halo Jumps to 192GB While Qualcomm Pushes On-Device AI AgentsAMD bumps its NPU mini PC to 192GB of unified memory, Qualcomm ships five agentic apps for Snapdragon X, and Korea rethinks its NPU strategy.
  43. 100 viewsA Robot Arm Joins the NPU Party: FastFlowLM Adds a Vision-Language-Action Model to Ryzen AIFastFlowLM's first stable release puts a robotics policy on Ryzen AI's NPU, while Korea ships a boxed NPU appliance and a Raspberry Pi learns to narrate what it sees.
  44. 116 viewsKorea's KT Ships a Boxed NPU LLM Station While the ESP32 Crowd Trims Memory FurtherA Korean telecom sells an all-in-one on-prem LLM box built on a domestic NPU, while ESP32 tinkerers keep shrinking what a model needs to run.
  45. 108 viewsRyzen AI's NPU Runtime Goes Official While a Raspberry Pi Learns to See and Speak with a Tiny LLMAMD folds a hobbyist NPU runtime into ROCm, Google shows Gemma driving a robot from a Raspberry Pi 5, and a 45M-parameter model books tool calls on a phone.
  46. 91 viewsRaspberry Pi's GPU Joins the AI Party While a 14MB Model Learns to Call ToolsA Raspberry Pi 5 runs Gemma and vision models split across CPU and GPU, and a 45M-parameter model fits tool-calling into 28MB of RAM.
  47. 99 viewsThe $8 AI Chip Grows a Coffee Habit: ESP32 Tiny-LLM Trick Gets a Second ModelA new 'Barista' model and a closer look at Google's Per-Layer Embeddings trick show what running an LLM on a microcontroller can and can't do.