Tag

llama.cpp

7 posts

ExecuTorch Hits 1.0 as llama.cpp Learns to Make Instant AI Decisions
Edge AI4 min read

ExecuTorch Hits 1.0 as llama.cpp Learns to Make Instant AI Decisions

PyTorch declares its on-device AI runtime production-ready while llama.cpp adds a decision-model endpoint and two new NPU modules land for edge boards.

89 views
Read
A $300 GPU Streams a 177B AI Model From an SSD While llama.cpp Learns to Skip Ahead
Edge AI4 min read

A $300 GPU Streams a 177B AI Model From an SSD While llama.cpp Learns to Skip Ahead

Community builders push token throughput further this week — via SSD-streamed MoE experts, prompt-lookup drafting, and a wrapper for Apple's built-in on-device LLM.

165 views
Read
llama.cpp's Nightly Grind Teaches Phone Chips New AI Tricks While NVIDIA Ships a Rival Edge Runtime
Edge AI4 min read

llama.cpp's Nightly Grind Teaches Phone Chips New AI Tricks While NVIDIA Ships a Rival Edge Runtime

Small llama.cpp builds keep adding Hexagon DSP ops, NVIDIA's TensorRT-Edge-LLM adds Day-0 model support, and a new leaderboard measures tokens per joule.

93 views
Read
Runtimes on the Move: llama.cpp, ExecuTorch and LiteRT All Update as Edge AI's Software Layer Speeds Up
Edge AI4 min read

Runtimes on the Move: llama.cpp, ExecuTorch and LiteRT All Update as Edge AI's Software Layer Speeds Up

llama.cpp shipped two builds in two days, ExecuTorch hit 1.0 with new NPU backends, and LiteRT tuned fp16 kernels for mobile CPUs.

135 views
Read
Runtimes Race Ahead: llama.cpp 0.4.0 and a 30B On-Device AI Agent Test the Edge's Limits
Edge AI4 min read

Runtimes Race Ahead: llama.cpp 0.4.0 and a 30B On-Device AI Agent Test the Edge's Limits

A new llama.cpp release, an ExecuTorch-powered 30B agent model, a cheap RK3576 vision board, and a DIY Jetson robot dog mark a busy week for edge toolchains.

101 views
Read
llama.cpp Teaches an NPU to Share the Load as Small AI Engines Keep Multiplying
Edge AI4 min read

llama.cpp Teaches an NPU to Share the Load as Small AI Engines Keep Multiplying

llama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.

117 views
Read
Edge Dispatch: Meta's Muse Glimmer Bets Big on On-Device Agentic AI as the Runtime Wars Keep Multiplying
Edge AI4 min read

Edge Dispatch: Meta's Muse Glimmer Bets Big on On-Device Agentic AI as the Runtime Wars Keep Multiplying

Meta ships an on-device agentic model, an MoE engine claims 753B on one GPU, and researchers find 10 CVEs in a local inference engine.

293 views
Read