
ExecuTorch Hits 1.0 as llama.cpp Learns to Make Instant AI Decisions
PyTorch declares its on-device AI runtime production-ready while llama.cpp adds a decision-model endpoint and two new NPU modules land for edge boards.
Tag
7 posts

PyTorch declares its on-device AI runtime production-ready while llama.cpp adds a decision-model endpoint and two new NPU modules land for edge boards.

Community builders push token throughput further this week — via SSD-streamed MoE experts, prompt-lookup drafting, and a wrapper for Apple's built-in on-device LLM.

Small llama.cpp builds keep adding Hexagon DSP ops, NVIDIA's TensorRT-Edge-LLM adds Day-0 model support, and a new leaderboard measures tokens per joule.

llama.cpp shipped two builds in two days, ExecuTorch hit 1.0 with new NPU backends, and LiteRT tuned fp16 kernels for mobile CPUs.

A new llama.cpp release, an ExecuTorch-powered 30B agent model, a cheap RK3576 vision board, and a DIY Jetson robot dog mark a busy week for edge toolchains.

llama.cpp adds multi-NPU Hexagon support, ONNX Runtime brings quantized KV caches to the browser, and a solo Rust engine beats llama.cpp on tiny models.

Meta ships an on-device agentic model, an MoE engine claims 753B on one GPU, and researchers find 10 CVEs in a local inference engine.