Back to Blog

pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Prateek SinghSeptember 26, 202610 min read59 views
pankhllm: The LLM Gateway That Learns to Skip the LLM, Without Replacing the Stack You Already Run

Most of an agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which parameters, made thousands of times a day by a model paid in seconds and tokens. pankhllm sits where your app already calls an LLM, learns those decisions from its own traffic, and starts making them in 0.2 ms on a CPU with a 262 KB model. What it is not sure about still goes to your LLM. On the same 14 questions: a 12B planner 1,743 ms, Laya 49 ms, pankhllm's own model 4 ms, all 14 correct. Here is what it is, what it is not, and where it stops.

Most of your agent's LLM calls are not writing anything. They are decisions: which tool, which skill, which report, which parameters. Your application pays a large model seconds and tokens to make the same decision thousands of times a day, and it makes the same decision every time.

pankhllm started as an internal fix for exactly that: agents in production paying LLM latency for decisions they had already made many times. It sits where your app already calls an LLM, watches those decisions, trains its own tiny model on them, and starts making them itself in 0.2 ms on a CPU. What it is not sure about still goes to the LLM you already use, unchanged, and becomes tomorrow's training data. Nothing in the stack is replaced. One Rust binary, OpenAI-compatible, MIT, public as of today.

same 14 questions, median answer
1,743 ms → 4 ms
a 12B planner against pankhllm's own model; 14 of 14 correct both ways
one decision, in-process, one CPU core
0.2 ms
no GPU, no LLM call
every decision model together
262 KB
a hashed word-and-bigram logistic regression: 12,545 weights for a seven-operation catalog
2,000 unseen pharma KPI questions
63%, 0 wrong
answered with no LLM call; the other 37% went to the LLM as before

All from runs in the repository: the 14-question race on an RTX 4090 Laptop (planner gemma4:12b via Ollama); pharma-KPI tables CPU-only; timings on an Intel i9-13900HK, one core. docs/BENCHMARK-DECISIONS.md, docs/TRAINING.md.

from openai import OpenAI
client = OpenAI(base_url="http://localhost:4000/v1", api_key="...")   # was: your LLM provider
# the same one-line change in LangChain, Semantic Kernel, the Vercel AI SDK, or any OpenAI client

Not another router: where it sits

This is not an empty field. LiteLLM owns much of the "LLM gateway" idea; Laya is the closest thing technically, close enough that I use it as a lane inside pankhllm and in another project. In their own words:

what it fundamentally doeswhat answers a decision
LiteLLMA gateway across 100+ LLM APIs in the OpenAI format: routing, retries, budgets, guardrails, logging. It chooses which LLM to call.an LLM
LayaA 421M-parameter decision model: typed yes/no, choice and score in one forward pass, about 33 ms on a T4. A small neural model instead of an LLM.a 421M transformer
pankhllmA gateway that records the decisions its own traffic already made, trains a 262 KB model on them, and makes the familiar ones itself. It removes inference from decisions the application has already learned.12,545 weights, or nothing
LiteLLM from its repository description; Laya from its README (three checkpoints: ModernBERT-large 421M for English and typed decisions, mmBERT-base 322M multilingual); pankhllm from its README.

The last column is the whole positioning. pankhllm asks: once my application has seen this decision enough times, why should any model run at all? Its model is not a transformer. It is a hashed word-and-bigram logistic regression over the shape of the question, 12,545 weights for a seven-operation catalog, 262 KB in total, trained by the router itself in about a second on one CPU core.

That makes it a cascade layer, not a rival router. RouteLLM and Not Diamond answer "which LLM should handle this request?", and answer it well; LiteLLM carries that call to a provider. pankhllm asks the question one step earlier: does an LLM need to handle this decision at all? Put LiteLLM underneath it and nothing conflicts. Put Laya in the middle and the questions pankhllm declines get a 33 ms answer before anything reaches a generative model. The cheapest LLM call is the one your application learns not to make.

The measured continuum

The repository's first benchmark is that ladder on one machine: fourteen labelled questions, ten answerable by a three-operation catalog and four that must go to the agent, same executor, one variable changed.

Same machine, same 14 questions, same executor; one variable: how the router decides log scale · wide bar: median end to end · thin bar: the decision step alone · all three answered 14 of 14 0.1 ms1 ms10 ms100 ms1 s10 s Generative planner gemma4:12b via Ollama, RTX 4090 Laptop 1,743 ms end to end decision alone: ~1.4 s Laya, decision model 0.3.20, English checkpoint, same GPU 49 ms end to end decision alone: 12 ms pankhllm's own model self-trained on 117 planner labels, CPU 4 ms end to end decision alone: 0.13 to 0.26 ms
From docs/BENCHMARK-DECISIONS.md: laptop with an RTX 4090 Laptop GPU (16 GB); planner and agent gemma4:12b on local Ollama, reasoning off; Laya 0.3.20 via laya-serve, English checkpoint, on the same GPU; pankhllm's own model trained from 117 planner labels on questions that did not overlap the test set.

Two things matter more than the 4 ms. Laya got 14 of 14 too, at 49 ms; a 421M decision model is a very good middle lane. And the 4 ms row is not "we beat Laya"; it is a decision seen 117 times, so no model runs. Zero-shot over the whole catalog, Laya picked right on 8 of 14; constraining it to the operations the question structurally fills is what got it to 14, and that constraint is the same one pankhllm's own lane uses.

How it learns from its own traffic

lanewhat answerslatencyLLM calls
Cachean identical question, same tenant and data version~1 ms0
Learned routea question shape seen before, new entities or phrasing~1 ms + your tool0
Rulea pattern you pinned~1 ms + your tool0
Decisionpankhllm's own model picks the operation and fills the parameters0.2 ms + your tool0
Plannera new question that maps to an operation in your catalogone small-model call + your executor1
Agentexplanation, judgement, open-ended workbest eligible model, hedgedas needed
The six lanes, from the README. Every lane fails open to the next, so the worst case is what you had before.

Two lanes above the decision model do more than caching. Learned routes keep a question's shape with its argument template: "calls for T-123 in 2026-08" becomes "calls for {territory} in {period}" with get_calls(territory, period), and the next territory and month never need the original decision; only flat typed arguments are learned, never code or SQL, and only from two clean tool calls. Speculative answer templates have the model write the answer's template on the tool-selection hop, so a learned two-hop turn completes with zero model calls.

The decision lane is the new part. The honest name for the loop is application-level distillation: not compressing a model's weights, but distilling the repetitive decisions it makes inside one application into something that needs no model. The miner retrains, learns routes the live path missed, quarantines routes that fail twice, and recommends the threshold for your precision target.

# observe: every clean planner decision and tool call becomes a label; only the question's SHAPE is kept
# train: hold out a fifth of the shapes, pick the threshold for your precision target
pankhllm mine --apply                          # ~1.1 s for 20,000 questions, one CPU core, 45 MB peak
# decide: act only among operations the question fills exactly, only when confident and the wording is familiar
# or teach it up front from your skills, prompts and logs, with any teacher:
python3 notebooks/run_trainer.py --skills ./skills --prompts agent.yaml --logs app.jsonl --teacher ollama

The precision comes from what surrounds the model. A slot filler first finds the operations the question fills exactly, from every code, date and number in it; the model chooses only among those plus "unsupported"; a vote the structure disagrees with, or unfamiliar wording, goes to the planner. Two candidates for a slot, or none, is a planner question, never a guess. The router never builds a query; a validated plan is POSTed to your executor. ML for fuzzy intent, deterministic constraints for correctness, an LLM for ambiguity.

Where it is weaker, from the same document

A bag-of-words model has a ceiling on unseen wording, and the repository measured where it is: seven operations plus a reject class, a fictional pharma-KPI catalog, CPU only, held-out test sets.

model trained ontest setdecided without a modelprecision
+ 183 teacher-labelled questions2,000 held-out templates63.3%100% (0 wrong of 1,266)
+ 183 teacher-labelled questions82 model-written questions67.1%100%
+ 183 teacher-labelled questions184 model-written, held-out half61.4%96.5% (4 wrong of 113)
From docs/BENCHMARK-DECISIONS.md, "Pharma KPI catalog at scale": 20,000 template questions plus 183 teacher-labelled ones. Every question not decided went to the generative planner.

Trained on templates alone, the model was 100% on template phrasing and 96% on model-written questions; templates do not generalise to how people write. 183 teacher labels closed most of the gap, and in production they arrive free, because every planner decision is one. The disagreement check trades coverage for precision on purpose: before it, the harder set was 65% decided at 93.7%. The four disagreements in the last row, read against the teacher: one the model got right, two defensible either way, one wrong. That is 96.5% at the row level.

One thing from the journal I would rather say than bury: label --out was found recording already-labelled rows into the training store, so test rows could reach training. It was fixed, a binary-level test guards it, and the numbers above are from after the fix.

The limits the numbers do not show are the real ones. The evidence is self-generated: template questions and teacher-written questions; real human traffic from another organisation does not exist yet. Seven operations is a small catalog; 30, 100 or 500 is unmeasured. Drift: an operation whose meaning changes will have taught a stale decision; quarantine helps, versioned schemas and label provenance do not exist yet. Poisoned learning: traffic that becomes labels is a feedback loop; recording only clean executions, plus shadow mode, is the floor, not the ceiling.

So judge it the way a production system is judged, not a router benchmark: p50, p95 and p99 before and after; the share of requests that bypass the LLM; the precision of those bypassed requests; model calls eliminated; what happens under drift and failure; and how fast coverage climbs from real traffic. One genuine deployment where 40 to 60% of recurring decisions move from seconds to sub-millisecond overhead while the uncertain ones stay untouched would be worth more than every table above.

Try it

It is on crates.io, PyPI and npm. The 0.1.1 release, one version number everywhere plus server wheels for macOS, Linux arm64 and Windows, is publishing as this goes out; until your wheel lands, cargo install pankhllm builds from source.

cargo install pankhllm                         # the router, from crates.io; pip install pankhllm for the Python client
export OPENAI_API_KEY=...                      # or ANTHROPIC_API_KEY, AZURE_OPENAI_API_KEY, none for Ollama
pankhllm check                                 # validate config, list models and prompts
pankhllm route "Why did the migration fail?"   # explain a routing decision, no model call
pankhllm serve --port 4000                     # docker build -t pankhllm . and cargo build --release also work

Change one base URL and read pankhllm.model in every response: the lane and operation that answered, with reasons. A worked agent setup is in docs/RECIPE-AGENTS.md; training from your own skills and logs in docs/TRAINING.md.

pankhllm is MIT, 117 offline tests, and a day old in public. It is not an LLM and does not replace yours. The graph that would prove the thesis is the one I do not have yet: the fraction of an agent's decisions compiled out of the generative path, day by day on real traffic, with the precision of the skipped calls beside it. If you run it in shadow mode, the agreement numbers it prints are the first points on that graph.

FAQ: pankhllm, LiteLLM and Laya

Is pankhllm a LiteLLM alternative?

No; it sits above one. LiteLLM chooses which LLM to call; pankhllm removes the calls that never needed a model and hands the rest to LiteLLM or a provider. Both speak the OpenAI wire format.

How is pankhllm different from Laya?

Laya is a 421M-parameter decision model: one forward pass, about 33 ms on a T4, options supplied at request time. pankhllm's model is a 262 KB logistic regression trained on the application's own traffic, fixed label space, 0.2 ms on a CPU; it can use Laya as the next lane for what it declines.

What happens when pankhllm is not sure?

The question goes to the planner or the agent as before, and that decision becomes a training label. The model acts only among operations the question structurally fills, above a threshold the miner picks for your precision target, and only on familiar wording.

How accurate is a 262 KB model?

On a seven-operation pharma-KPI catalog: 63% of 2,000 unseen template questions decided with 0 wrong, 67% of 82 model-written with 0 wrong, 61% of a harder set at 96.5%. Everything undecided went to the LLM. The precision comes from the structural constraint as much as the model.

Does pankhllm need a GPU?

No. Training on 20,000 questions takes 1.1 seconds on one CPU core and 45 MB; a decision takes 0.2 ms; the server is an 11 MB binary at 17 MB idle. The GPU in the benchmark ran the LLM planner and Laya, not pankhllm.

References & Citations

  • pankhllm README and docs/BENCHMARK-DECISIONS.md (the 14-question three-setup run: laptop, RTX 4090 Laptop GPU 16 GB, gemma4:12b via Ollama, Laya 0.3.20 via laya-serve; the pharma-KPI tables, CPU only, release build), docs/TRAINING.md "Resource needs" (Intel i9-13900HK, one core), docs/JOURNAL.md; singhpratech.github.io/pankhllm. Repository public 26 September 2026; crates.io 0.1.0, npm 0.1.0 and PyPI (pankhllm, pankhllm-server 0.1.1) published the same afternoon; the npm client, NuGet and the GitHub release binaries pending at the time of writing.
  • LiteLLM, BerriAI: described from the repository's own description ("Open Source AI Gateway for 100+ LLMs… Call any LLM in OpenAI format").
  • Laya, Convai Innovations, Apache 2.0: README checkpoint table (ModernBERT-large 421M English and typed-decisions, mmBERT-base 322M multilingual), 33 ms single decision on a T4, Router preload.

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article

Related Posts

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
Data Engineering19 min read

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Every GPU data library benchmarks itself at 50 million rows. Your dataframe has 80,000. On an Apple M4 Max, summing 1,000 integers takes the GPU 112 microseconds and Polars less than one: the GPU is more than 100 times behind. I built one of these libraries, so I measured the row count where the GPU overtakes the fastest CPU code for 111 operations: the median needs 10,000,000 rows against a multi-core library, about a million against one core, and sixteen never get there. So ArrowMetal 0.2.0 refuses the GPU below the line, byte-identical, and around that router it grew GPU readers for CSV, JSON, nested Parquet, Delta Lake and Iceberg, a Polars engine and a DuckDB optimizer extension.

58 views
Read
DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower
Data Engineering15 min read

DuckDB on the Apple Silicon GPU: Plain SQL, 19 of 22 TPC-H Queries on the Mac's Own GPU, and a Rule That Says Never Slower

DuckDB has no GPU backend of its own, and the GPU engines built for it need an NVIDIA card. gpudb 0.7 is my Apache-2.0 DuckDB extension for the GPU already inside your Mac, and for CUDA too. You write plain DuckDB SQL; the GPU takes a statement only where it has been measured faster than DuckDB on your own machine. On an Apple M4 Max, 19 of 22 TPC-H SF10 queries run on the Metal GPU at 1.06x to 48x with zero rows differing. The one row below parity is printed, not dropped.

78 views
Read
Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway
Data Engineering15 min read

Apple's GPU Has No 64-bit Floats. I Made It Sort 50 Million Doubles Anyway

The Metal Shading Language has no double type, and float64 is the default number in Python, pandas and Apache Arrow. Building ArrowMetal meant getting past three walls: a GPU with no 64-bit floats, a missing 64-bit atomic add, and a Swift compiler bug that reports errors nobody threw. Here is how each one was solved, what it cost, and why a GPU that cannot add two doubles sorts 50,000,000 of them in 32 ms.

118 views
Read