Tag

LLM

10 posts

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run
Data Engineering10 min read

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.

62 views
Read
When an AI-Generated Answer Actually Works, Who Is the Rule Protecting? A Fair Look at Both Sides — and a Rule Both Could Sign
Opinion10 min read

When an AI-Generated Answer Actually Works, Who Is the Rule Protecting? A Fair Look at Both Sides — and a Rule Both Could Sign

A programming community I take part in watches for AI-generated answers. The moderators have real reasons; so do the people the rule lands on. Here are both cases argued fairly, the three things both sides actually want, and a tiered guideline — verification, disclosure, capacity — that is stricter than a blanket ban, not looser.

135 views
Read
Flipper One Wants to Be the First Hacker Tool With a Local LLM — What Its 6 TOPS NPU Can Actually Run
Edge AI7 min read

Flipper One Wants to Be the First Hacker Tool With a Local LLM — What Its 6 TOPS NPU Can Actually Run

Flipper Devices' pocket Linux box promises an LLM that runs offline and knows the device inside out. Rockchip's own numbers say what a 6 TOPS RK3576 really does — and the NPU driver isn't in the kernel Flipper chose.

173 views
Read
AMD Just Bought the 'Ollama of NPUs': What FastFlowLM Means for Local LLMs
Edge AI6 min read

AMD Just Bought the 'Ollama of NPUs': What FastFlowLM Means for Local LLMs

A 17MB runtime that runs LLMs on AMD's Ryzen AI NPUs — built by three academics, acquired by AMD on July 17, 2026, folded into ROCm in August. The NPU rung of the edge ladder just got real.

227 views
Read
The Edge AI Compute Ladder: What Actually Runs on Every Board From $5 to $600
Edge AI8 min read

The Edge AI Compute Ladder: What Actually Runs on Every Board From $5 to $600

A $5 Pico 2 writes TinyStories. A $15 Pi Zero 2 W runs SmolLM2-135M. A $299 RISC-V board claims 30B. What AI really fits at every rung of the edge hardware ladder.

231 views
Read
The 180M-Parameter LLM Running on a $10 Microcontroller — and Almost Nobody Noticed
Edge AI7 min read

The 180M-Parameter LLM Running on a $10 Microcontroller — and Almost Nobody Noticed

On August 5, 2026 a 180.9M-parameter mixture-of-experts LLM ran on a $6-10 ESP32-P4 — and got two Hacker News points. The undercovered microcontroller AI story of the year.

113 views
Read
Running an LLM on an $8 Microcontroller: What's Real in 2026
Edge AI9 min read

Running an LLM on an $8 Microcontroller: What's Real in 2026

In July 2026 a 28.9-million-parameter LLM ran fully on-device on an $8 ESP32-S3 at almost 10 tokens per second. Here's how the trick works, who built it, what's hype, and what you can actually build with a microcontroller LLM today.

226 views
Read
Speculative Decoding: The Inference Trick Hiding in Plain Sight
Research7 min read

Speculative Decoding: The Inference Trick Hiding in Plain Sight

Speculative decoding promises faster LLM inference without touching model weights. The math holds up — but the production story is messier than the papers admit.

104 views
Read
RAG Is Dead, Long Live Agentic RAG: The Evolution of AI Knowledge Systems
AI & Machine Learning9 min read

RAG Is Dead, Long Live Agentic RAG: The Evolution of AI Knowledge Systems

Traditional RAG retrieves documents and stuffs them into context. Agentic RAG plans queries, evaluates results, and iterates until it finds the right answer.

76 views
Read
The Agentic Paradigm Shift: Why 2025 Changed Everything in AI Development
AI & Machine Learning9 min read

The Agentic Paradigm Shift: Why 2025 Changed Everything in AI Development

The shift from AI-as-tool to AI-as-agent represents the biggest paradigm change since the internet. Here's how we got here and where it's heading.

74 views
Read