pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.
Most of your agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which report, which parameters. Your application pays a large model seconds and tokens to make the same decision thousands of times a day, and it makes the same decision every time.
pankhllm started as an internal fix for exactly that: agents in production paying LLM latency for decisions they had already made many times. It sits where your app already calls an LLM, watches those decisions, trains its own tiny model on them, and starts making them itself in 0.2 ms on a CPU. What it is not sure about still goes to the LLM you already use, unchanged, and becomes tomorrow's training data. Nothing in the stack is replaced. One Rust binary, OpenAI-compatible, MIT, public as of today.
All from runs in the repository: the 14-question race on an RTX 4090 Laptop (planner gemma4:12b via Ollama); pharma-KPI tables CPU-only; timings on an Intel i9-13900HK, one core. docs/BENCHMARK-DECISIONS.md, docs/TRAINING.md.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:4000/v1", api_key="...") # was: your LLM provider
# the same one-line change in LangChain, Semantic Kernel, the Vercel AI SDK, or any OpenAI client
Not another router: where it sits
This is not an empty field. LiteLLM owns much of the "LLM gateway" idea; Laya is the closest thing technically, close enough that I use it as a lane inside pankhllm and in another project. In their own words:
| what it fundamentally does | what answers a decision | |
|---|---|---|
| LiteLLM | A gateway across 100+ LLM APIs in the OpenAI format: routing, retries, budgets, guardrails, logging. It chooses which LLM to call. | an LLM |
| Laya | A 421M-parameter decision model: typed yes/no, choice and score in one forward pass, about 33 ms on a T4. A small neural model instead of an LLM. | a 421M transformer |
| pankhllm | A gateway that records the decisions its own traffic already made, trains a 262 KB model on them, and makes the familiar ones itself. It removes inference from decisions the application has already learned. | 12,545 weights, or nothing |
The last column is the whole positioning. pankhllm asks: once my application has seen this decision enough times, why should any model run at all? Its model is not a transformer. It is a hashed word-and-bigram logistic regression over the shape of the question, 12,545 weights for a seven-operation catalog, 262 KB in total, trained by the router itself in about a second on one CPU core.
That makes it a cascade layer, not a rival router. RouteLLM and Not Diamond answer "which LLM should handle this request?", and answer it well; LiteLLM carries that call to a provider. pankhllm asks the question one step earlier: does an LLM need to handle this decision at all? Put LiteLLM underneath it and nothing conflicts. Put Laya in the middle and the questions pankhllm declines get a 33 ms answer before anything reaches a generative model. The cheapest LLM call is the one your application learns not to make.
The measured continuum
The repository's first benchmark is that ladder on one machine: fourteen labelled questions, ten answerable by a three-operation catalog and four that must go to the agent, same executor, one variable changed.
Two things matter more than the 4 ms. Laya got 14 of 14 too, at 49 ms; a 421M decision model is a very good middle lane. And the 4 ms row is not "we beat Laya"; it is a decision seen 117 times, so no model runs. Zero-shot over the whole catalog, Laya picked right on 8 of 14; constraining it to the operations the question structurally fills is what got it to 14, and that constraint is the same one pankhllm's own lane uses.
How it learns from its own traffic
| lane | what answers | latency | LLM calls |
|---|---|---|---|
| Cache | an identical question, same tenant and data version | ~1 ms | 0 |
| Learned route | a question shape seen before, new entities or phrasing | ~1 ms + your tool | 0 |
| Rule | a pattern you pinned | ~1 ms + your tool | 0 |
| Decision | pankhllm's own model picks the operation and fills the parameters | 0.2 ms + your tool | 0 |
| Planner | a new question that maps to an operation in your catalog | one small-model call + your executor | 1 |
| Agent | explanation, judgement, open-ended work | best eligible model, hedged | as needed |
Two lanes above the decision model do more than caching. Learned routes keep a question's shape with its argument template: "calls for T-123 in 2026-08" becomes "calls for {territory} in {period}" with get_calls(territory, period), and the next territory and month never need the original decision; only flat typed arguments are learned, never code or SQL, and only from two clean tool calls. Speculative answer templates have the model write the answer's template on the tool-selection hop, so a learned two-hop turn completes with zero model calls.
The decision lane is the new part. The honest name for the loop is application-level distillation: not compressing a model's weights, but distilling the repetitive decisions it makes inside one application into something that needs no model. The miner retrains, learns routes the live path missed, quarantines routes that fail twice, and recommends the threshold for your precision target.
# observe: every clean planner decision and tool call becomes a label; only the question's SHAPE is kept
# train: hold out a fifth of the shapes, pick the threshold for your precision target
pankhllm mine --apply # ~1.1 s for 20,000 questions, one CPU core, 45 MB peak
# decide: act only among operations the question fills exactly, only when confident and the wording is familiar
# or teach it up front from your skills, prompts and logs, with any teacher:
python3 notebooks/run_trainer.py --skills ./skills --prompts agent.yaml --logs app.jsonl --teacher ollama
The precision comes from what surrounds the model. A slot filler first finds the operations the question fills exactly, from every code, date and number in it; the model chooses only among those plus "unsupported"; a vote the structure disagrees with, or unfamiliar wording, goes to the planner. Two candidates for a slot, or none, is a planner question, never a guess. The router never builds a query; a validated plan is POSTed to your executor. ML for fuzzy intent, deterministic constraints for correctness, an LLM for ambiguity.
Where it is weaker, from the same document
A bag-of-words model has a ceiling on unseen wording, and the repository measured where it is: seven operations plus a reject class, a fictional pharma-KPI catalog, CPU only, held-out test sets.
| model trained on | test set | decided without a model | precision |
|---|---|---|---|
| + 183 teacher-labelled questions | 2,000 held-out templates | 63.3% | 100% (0 wrong of 1,266) |
| + 183 teacher-labelled questions | 82 model-written questions | 67.1% | 100% |
| + 183 teacher-labelled questions | 184 model-written, held-out half | 61.4% | 96.5% (4 wrong of 113) |
Trained on templates alone, the model was 100% on template phrasing and 96% on model-written questions; templates do not generalise to how people write. 183 teacher labels closed most of the gap, and in production they arrive free, because every planner decision is one. The disagreement check trades coverage for precision on purpose: before it, the harder set was 65% decided at 93.7%. The four disagreements in the last row, read against the teacher: one the model got right, two defensible either way, one wrong. That is 96.5% at the row level.
One thing from the journal I would rather say than bury: label --out was found recording already-labelled rows into the training store, so test rows could reach training. It was fixed, a binary-level test guards it, and the numbers above are from after the fix.
The limits the numbers do not show are the real ones. The evidence is self-generated: template questions and teacher-written questions; real human traffic from another organisation does not exist yet. Seven operations is a small catalog; 30, 100 or 500 is unmeasured. Drift: an operation whose meaning changes will have taught a stale decision; quarantine helps, versioned schemas and label provenance do not exist yet. Poisoned learning: traffic that becomes labels is a feedback loop; recording only clean executions, plus shadow mode, is the floor, not the ceiling.
So judge it the way a production system is judged, not a router benchmark: p50, p95 and p99 before and after; the share of requests that bypass the LLM; the precision of those bypassed requests; model calls eliminated; what happens under drift and failure; and how fast coverage climbs from real traffic. One genuine deployment where 40 to 60% of recurring decisions move from seconds to sub-millisecond overhead while the uncertain ones stay untouched would be worth more than every table above.
Try it
It is on crates.io, PyPI and npm. The 0.1.1 release, one version number everywhere plus server wheels for macOS, Linux arm64 and Windows, is publishing as this goes out; until your wheel lands, cargo install pankhllm builds from source.
cargo install pankhllm # the router, from crates.io; pip install pankhllm for the Python client
export OPENAI_API_KEY=... # or ANTHROPIC_API_KEY, AZURE_OPENAI_API_KEY, none for Ollama
pankhllm check # validate config, list models and prompts
pankhllm route "Why did the migration fail?" # explain a routing decision, no model call
pankhllm serve --port 4000 # docker build -t pankhllm . and cargo build --release also work
Change one base URL and read pankhllm.model in every response: the lane and operation that answered, with reasons. A worked agent setup is in docs/RECIPE-AGENTS.md; training from your own skills and logs in docs/TRAINING.md.
pankhllm is MIT, 117 offline tests, and a day old in public. It is not an LLM and does not replace yours. The graph that would prove the thesis is the one I do not have yet: the fraction of an agent's decisions compiled out of the generative path, day by day on real traffic, with the precision of the skipped calls beside it. If you run it in shadow mode, the agreement numbers it prints are the first points on that graph.
FAQ: pankhllm, LiteLLM and Laya
Is pankhllm a LiteLLM alternative?
No; it sits above one. LiteLLM chooses which LLM to call; pankhllm removes the calls that never needed a model and hands the rest to LiteLLM or a provider. Both speak the OpenAI wire format.
How is pankhllm different from Laya?
Laya is a 421M-parameter decision model: one forward pass, about 33 ms on a T4, options supplied at request time. pankhllm's model is a 262 KB logistic regression trained on the application's own traffic, fixed label space, 0.2 ms on a CPU; it can use Laya as the next lane for what it declines.
What happens when pankhllm is not sure?
The question goes to the planner or the agent as before, and that decision becomes a training label. The model acts only among operations the question structurally fills, above a threshold the miner picks for your precision target, and only on familiar wording.
How accurate is a 262 KB model?
On a seven-operation pharma-KPI catalog: 63% of 2,000 unseen template questions decided with 0 wrong, 67% of 82 model-written with 0 wrong, 61% of a harder set at 96.5%. Everything undecided went to the LLM. The precision comes from the structural constraint as much as the model.
Does pankhllm need a GPU?
No. Training on 20,000 questions takes 1.1 seconds on one CPU core and 45 MB; a decision takes 0.2 ms; the server is an 11 MB binary at 17 MB idle. The GPU in the benchmark ran the LLM planner and Laya, not pankhllm.
Read next
References & Citations
- pankhllm README and docs/BENCHMARK-DECISIONS.md (the 14-question three-setup run: laptop, RTX 4090 Laptop GPU 16 GB, gemma4:12b via Ollama, Laya 0.3.20 via laya-serve; the pharma-KPI tables, CPU only, release build), docs/TRAINING.md "Resource needs" (Intel i9-13900HK, one core), docs/JOURNAL.md; singhpratech.github.io/pankhllm. Repository public 26 September 2026; crates.io 0.1.0, npm 0.1.0 and PyPI (pankhllm, pankhllm-server 0.1.1) published the same afternoon; the npm client, NuGet and the GitHub release binaries pending at the time of writing.
- LiteLLM, BerriAI: described from the repository's own description ("Open Source AI Gateway for 100+ LLMs… Call any LLM in OpenAI format").
- Laya, Convai Innovations, Apache 2.0: README checkpoint table (ModernBERT-large 421M English and typed-decisions, mmBERT-base 322M multilingual), 33 ms single decision on a T4, Router preload.
Subscribe to new posts from theaivibe.org
Related Posts
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.
DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower
DuckDB has no GPU backend of its own, and the GPU engines built for it need an NVIDIA card. gpudb 0.7 is my Apache-2.0 DuckDB extension for the GPU already inside your Mac, and for CUDA too. You write plain DuckDB SQL; the GPU takes a statement only where it has been measured faster than DuckDB on your own machine. On an Apple M4 Max, 19 of 22 TPC-H SF10 queries run on the Metal GPU at 1.06x to 48x with zero rows differing. The one row below parity is printed, not dropped.
Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway
The Metal Shading Language has no double type, and float64 is the default number in Python, pandas and Apache Arrow. Building ArrowMetal meant getting past three walls: a GPU with no 64-bit floats, a missing 64-bit atomic add, and a Swift compiler bug that reports errors nobody threw. Here is how each one was solved, what it cost, and why a GPU that cannot add two doubles sorts 50,000,000 of them in 32 ms.