Edge AI4 min read
Qualcomm Chases 30B AI Models on a Phone as Edge Silicon Keeps Multiplying
A phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.
32 views
Read Tag
3 posts
A phone NPU claims 30B MoE inference, a 35B model streams from storage on a Mac, and an XDNA1 NPU gets a Linux bring-up.

Fresh arXiv work rethinks quantization strategy and sustainability, an Apple-adjacent paper shrinks the dictation encoder, and new silicon and Jetson guidance round out the ed

Three chipmakers and a mini-PC builder all attack the same problem this week: getting AI inference closer to memory, not just closer to silicon.