Does It Earn a Lane? Quantize, Measure, and Decide

Lesson 10: int8 quantization of the Stage C deciders, CPU latency and size, a stdlib server speaking pankhllm's decision shape, and a real run through pankhllm's own binary. The default int8 broke Qwen; the small decider got smaller and slower. And the contract, applied honestly: neither lane earns its place yet.
A model that decides well in a notebook still has to earn a place in production: fast, small, and right often enough that the slow path is not needed. Neither of mine does yet. This last Stage C lesson of e=mc² serves the Lesson 09 deciders, in int8, the way pankhllm talks to external decision engines. On teacher-written questions, with the known-words gate on, the int8 decider is more precise than pankhllm's regression but decides less, and neither reaches 98%.
Where you are: Lesson 09 left you pankh-decider-v2 (27.9% of teacher-written questions at 95.8%) and a Qwen adapter (65.1% at 88.4%), neither at 98%. This last lesson shrinks both to int8, measures milliseconds and megabytes, and lets pankhllm's own binary judge whether either earns a lane, its word for a decision engine it trusts first.
One run, 11 October 2026, Apple Silicon Mac CPU, pankhllm 0.1.1. The whole notebook took 181.5 seconds; not timed on Colab.
The house rule, and a smaller backpack
Momo can dispatch, fill in forms and has studied the teacher's decisions. Before he gets his own desk, a lane in pankhllm's words, Ella (tonight she sets the house rule and asks Momo through /v1/systemone) lays it down: act alone only if, on questions you never saw, you are right 98 times in 100. And travel light: his notes go into a smaller backpack, written with fewer digits.
The backpack is int8 dynamic quantization: each weight of a linear layer becomes a one-byte whole number times a scale; activations are scaled on the fly. The rule is pankhllm's: its miner targets 98% precision, and a lane earns its place only if it holds that on unseen questions and decides more than the existing 262 KB regression.
| Word | In Momo's world | Grown-up meaning |
|---|---|---|
| fp32, int8 | Dials written with seven digits; the same dials rounded to whole numbers | 4-byte floats; 1-byte integers times a scale |
| Per-tensor, per-channel | One scale for the page; one per line | One scale per weight matrix; one per output row |
| Dynamic quantization | Ella rounding what Momo hears as he hears it | Weights int8 ahead of time, activations quantized at run time |
| Outlier | One shout that drowns every whisper | A few huge activation values that crush the rest under one shared scale |
| Lane, contract | Momo's own desk; the house rule of 98 in 100 | A decision engine pankhllm consults first; precision of at least 98% plus more coverage than today |
| /v1/systemone | The window where Ella asks Momo a typed question | pankhllm's HTTP shape for external decision engines such as Laya, its published GPU decision model, which this decider is meant to be a CPU-sized cousin of |
| Shadow mode | Momo answers on paper while Ella still decides | Run the lane beside the planner and measure before acting |
Where the picture is wrong: integer weights are prepared once, not rounded at serve time.
Quantizing, three times
PyTorch's one-liner, quantize_dynamic on every linear layer, worked on the 0.84M-parameter decider: decisions unchanged on 99.0% of the 300-question check and 99.59% over all 2,172 test questions, not identical. On Qwen2.5-0.5B, adapter merged, it kept 44.7%. Per-channel scales, one per output row instead of one per matrix, with the output head in fp32, kept 48.7%. Measuring every linear layer's input on 200 calibration questions found outliers: a few feed-forward inputs near 1,700, as the LLM.int8() paper describes. Keeping those seven layers in fp32 brought agreement to 92.0%, and 94.29% over all test questions. Attempts two and three were chosen after seeing attempt one.
from torch.ao.quantization import quantize_dynamic, per_channel_dynamic_qconfig
attempt1 = quantize_dynamic(model, {nn.Linear}, dtype=torch.qint8) # the one-liner
keep = {"lm_head", *outlier_layers} # inputs above 100 on calibration questions
spec = {n: per_channel_dynamic_qconfig for n, m in model.named_modules()
if isinstance(m, nn.Linear) and n not in keep}
attempt3 = quantize_dynamic(model, spec, dtype=torch.qint8)
# what pankhllm sends the lane, and what the stdlib server answers
{"state": "who covers T-412?", "questions": {"operation": {"type": "choice",
"criteria": {"coverage_owner": "...", "prescriber_list": "...", "UNSUPPORTED": "..."}}}}
{"answers": {"operation": {"choice": "coverage_owner", "answer_confidence": 0.999851}}}
| Model | Teacher-written (172) | Held-out templates | CPU ms, p50 / p95 | MB |
|---|---|---|---|---|
| Decider v2, fp32 | 27.9% at 95.8% | 48.6% at 99.3% | 0.675 / 0.770 | 3.37 |
| Decider v2, int8 | 28.5% at 95.9% | 45.3% at 99.2% | 1.297 / 1.468 | 1.10 |
| Qwen2.5-0.5B LoRA, fp32 | 65.1% at 88.4% | 87.6% at 98.6% | 42.394 / 45.449 | 1,976.24 |
| Qwen2.5-0.5B LoRA, int8 (attempt 3) | 48.3% at 92.8% | 65.5% at 98.2% | 40.526 / 48.282 | 1,000.04 |
int8 cut the decider from 3.37 MB to 1.10 MB and made it slower: 1.30 ms against 0.68 ms in this run (0.64 in Lesson 09's) at the median, or p50 (p95, the time 95 answers in 100 beat, is in the table), because its matrices are too small for integer arithmetic to pay. Qwen halved rather than quartered, 1,976 MB to 1,000 MB, because its 545 MB embedding table stays fp32; speed was about the same.
Through pankhllm itself
pankhllm talks to external engines with one POST to /v1/systemone, defined in its decide.rs: the question plus a choice among the operations it structurally fits, UNSUPPORTED last. The notebook's HTTP server is plain Python standard library; a round trip took 1.62 ms. It also applies pankhllm's known-words gate: when under 70% of a question's words were seen in training, it gives no answer and pankhllm falls back.
The pankhllm 0.1.1 binary's decide-eval, its offline scoring command, runs labelled questions through its own slot filler, structural checks and decision engine. Trained on the 20,000 templates alone, its regression decided 52.45% of held-out templates at 100%, 1,049 right and none wrong: the published 52.5%, reproduced. Then the same with Lesson 09's 183 labels, then the two lanes, each with the gate on and off.
| Decider | Held-out templates (2,000) | Teacher-written | 14 benchmark questions | Decision ms (p50) | MB |
|---|---|---|---|---|---|
| pankhllm regression + 183 of its own teacher labels (published) | 63.3% at 100% | A: 67.1% at 100%; B: 61.4% at 96.5% | 14/14 (a separate model) | 0.21 to 0.36 | 0.262 (all models) |
| Laya external decision model (published, GPU) | not published | not published | 14/14 (8 fast, 2 fell back) | 12 | not published |
| pankhllm regression, templates only (measured, gate on) | 52.45% at 100% | 29.65% at 90.2% | 6 decided, 0 wrong | 0.17 | in its store |
| pankhllm regression + Lesson 09's 183 labels (gate on or off) | 61.95% at 100% | 36.63% at 92.06% | 6 decided, 0 wrong | 0.17 | in its store |
| pankh-decider-v2 int8 lane (gate on or off) | 55.4% at 100% | 30.23% at 96.15% | 7 decided, 0 wrong | 1.73 | 1.10 |
| Qwen2.5-0.5B LoRA int8 lane (gate on) | not run | 36.05% at 93.55% | 7 decided, 1 wrong | 36.9 | 1,000 |
| Qwen2.5-0.5B LoRA int8 lane (gate off) | not run | 36.63% at 92.06% | 7 decided, 1 wrong | 37.3 | 1,000 |
The gate changed nothing for the regression or the v2 lane at these thresholds, and removed one wrong decision from the Qwen lane. Without it, the Qwen lane and the regression landed on the same counts, 58 right and 5 wrong. The server does not answer pankhllm's follow-ups about unnamed enum values; those fall back to the planner.
The verdict
pankh-decider-v2 int8 does not earn a lane: with the gate, 30.2% of teacher-written questions decided at 96.2%, against the regression's 36.6% at 92.1%: more precise, less coverage, short of 98%. Qwen2.5-0.5B int8 does not earn a lane: 36.0% at 93.5%. The gate does not flip either verdict.
What it means, honestly
Quantization is easy to do and easy to get wrong: check agreement with fp32 first, and expect no speed-up on a model this small. On held-out templates every decider measured on them holds 98%; on questions a different author wrote, none does, pankhllm's regression included. The lesson's lane config therefore starts in shadow mode, answering beside the planner without acting until real labels bring the precision.
Run it yourself
- Any CPU. Download lesson_10.ipynb. It loads Lesson 09's models (a token is needed while that repo is private) and fetches the pankhllm binary for macOS arm64 or Linux x64.
- What is kept. The int8 decider,
pankhllm-lane.yamland RESULTS.md. The int8 Qwen is too large for the repo; the notebook rebuilds it.
The journey, in one breath
You built a language model from random dials (01), read what its heads look at (02), found the size where it starts to recite (03), taught it a new character and measured the forgetting (04), taught it to take requests (05), did that again with 2% of the dials (06), gave a tiny one a real job (07), a form to fill (08), real phrasings to learn from (09), and a smaller backpack and a judge (10).
What you have now: ten lessons' worth of weights: tiny-storyteller, v2, instruct, three LoRA adapters (Lesson 06's, and Lesson 09's BERT-mini and Qwen), pankh-decider v0 to v2, an int8 decider and a lane config in shadow mode. What you do not have: a lane; nothing holds 98% on other authors' questions. That is the next journey.
Check yourself
The number to reproduce: the verdict. This run, gate on: v2 int8 lane 30.2% at 96.2% on teacher-written questions, below the regression's coverage and below 98%.
- Why does Qwen shrink much less than four times?
- Why recalibrate the threshold after quantization instead of copying it from fp32?
- Why does pankhllm start a new lane in shadow mode?
Questions people ask
Why did int8 make the small model slower?
Its matrices are 128 wide. Quantizing activations and converting back costs more than the integer multiply saves; int8 pays off on large, memory-bound matrices, where fetching the weights, not the arithmetic, is the slow part.
Why did the default int8 break Qwen?
Dynamic quantization gives each activation tensor one scale, and a few of Qwen's feed-forward inputs reach about 1,700: the outlier problem of the LLM.int8() paper. Keeping those seven layers in fp32 brought agreement back to 94.29%.
So what would earn a lane?
More and better real phrasings, a larger teacher-written calibration set, and shadow mode on real traffic. The number format was not the problem; precision on other authors' questions was.
References & Citations
- singhpratech/emc2-lesson-10-int8 on the Hugging Face Hub: lesson_10.ipynb, RESULTS.md, the int8 decider and pankhllm-lane.yaml. Run of 11 October 2026.
- pankhllm (MIT): release 0.1.1, docs/BENCHMARK-DECISIONS.md and docs/TRAINING.md (source of the published rows), src/decide.rs (the /v1/systemone shape). Laya is the external engine pankhllm's benchmark used.
- PyTorch quantization (the eager API used here prints a deprecation notice pointing to torchao). Dettmers, T. et al. (2022), LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.
- All measured numbers are from one run on an Apple Silicon Mac CPU (qnnpack engine; latency with four threads, batch 1).
Subscribe to new posts from theaivibe.org