Application-Level Distillation: Learn From a Bigger Model's Own Decisions

Lesson 09: a teacher model writes 420 realistic questions and labels them, 406 kept, for 49 cents; 183 join 20,000 templates. Three small models learn from them side by side: the from-scratch decider, BERT-mini with LoRA, and Qwen2.5-0.5B with LoRA. Coverage on teacher-written questions, before and after, with the precision that came with it.
Templates teach a small model the catalog, not how people talk. Distillation does: let a bigger model handle realistic questions, keep its decisions, train the small model on them. pankhllm's published benchmark found that 183 such labels (the 183 you met in Lesson 07), added to 20,000 template questions, raised its coverage by about ten points without losing precision. This lesson of e=mc² repeats that step with three small models, from the same split, with the same 183 labels and the same tests; how each samples its training data differs, as listed below.
Where you are: Lesson 08 left you pankh-decider-v1: every answer valid, 71.3% exactly right, 28.7% valid but wrong, all on template phrasings. Real people do not write in templates. Tonight a bigger model writes and labels 420 real questions, and three students learn from 183 of them: v1, and two borrowed models wearing Lesson 06's LoRA notes.
One run, 11 October 2026, Apple Silicon Mac; the whole notebook took 169.6 seconds. 172 teacher-written test questions, none sharing a question or a shape with the teacher training or calibration set; two share a shape with the template training set.
Ella becomes the teacher
Momo, the model, has practised on Ella's sentence patterns; real customers do not speak in patterns. So Ella, who already answers everything Momo passes her, writes 420 realistic questions and labels them, and Momo studies those decisions. That is application-level distillation: learning from a bigger model's decisions inside the application. Hinton's distillation copies the teacher's probabilities; this lesson copies its decisions. In pankhllm the teacher is the same slow planner undecided questions already go to, so the labels cost nothing extra.
Two classmates join Momo: Google's 4-layer BERT-mini, an encoder with 11.2M parameters, and Qwen2.5-0.5B-Instruct, a decoder, 495M parameters with its adapters. Both are fine-tuned with LoRA through PEFT, parameter-efficient fine-tuning: borrowed weights frozen, small add-on matrices learned. This is where Lesson 06 pays: the sticky notes go on a borrowed book instead of Momo's own. All three read the question's shape, Lesson 07's trick of replacing codes, months and numbers with markers.
| Word | In Momo's world | Grown-up meaning |
|---|---|---|
| Teacher model | The big, slow dispatcher Ella invites | A large model that writes and labels questions |
| Shape | Stickers over codes and dates (Lesson 07) | The question with codes, months and numbers replaced by markers |
| Distillation | Momo copying how the big dispatcher decides | Training a small model on a large model's outputs |
| Label validation | Throwing out forms the counter would reject | Dropping teacher labels that fail the catalog check |
| Encoder, decoder | A reader who sees the whole question; a writer whose next word is the answer | BERT-style bidirectional model; GPT-style left-to-right model |
| LoRA, PEFT | Sticky notes on a borrowed book | Frozen pretrained weights plus small trainable low-rank matrices; parameter-efficient fine-tuning |
| Disagreement check | If Momo's favourite counter cannot take the form, Ella decides | Abstain when the top operation is not structurally feasible |
| Coverage at 98% precision | How many he handles while almost never wrong | Decided share at a threshold holding 98% precision |
Where the picture is wrong: the teacher has its own mistakes, so precision on its questions means agreement with it, not truth.
The teacher set, made tonight
pankhllm's planner is a local 12B model and its teacher-written questions are not public, so I had Claude Sonnet 4.5 via OpenRouter stand in for that planner. It wrote 420 questions from the catalog's descriptions, not its regular expressions (pankhllm's DATASETS.md records a teacher copying a regex into its questions), then labelled each against the catalog without knowing what it was written for. Validation dropped 14 labels, mostly months written as words; 406 remained. It took 49 calls and $0.49 as reported by OpenRouter. After removing near-duplicates, 405: 183 for training, 50 for calibration, 172 for the teacher-written test.
What the notebook compares
All three face pankhllm's disagreement check: choose among all eight classes, abstain if the favourite operation does not structurally fit the question, otherwise decide when confidence clears that model's threshold, chosen as the lowest at which both calibration sets hold 98%. Lesson 08's renormalised confidence gave a lone feasible operation 1.0; this check does not.
- From scratch. Before is Lesson 08's v1. After continues it for 300 steps with 8 teacher questions per batch of 64, so each label is seen about 13 times: pankh-decider-v2.
- BERT-mini with LoRA. Rank 8 on two attention matrices per layer, a classification head (one small layer that turns the encoder's summary into a choice), 5,892 template shapes plus the labels three times, 2 epochs (an epoch is one pass over all the training examples): about 6 views per label.
- Qwen with LoRA. Rank 8 on four attention projections (query, key, value and output); the next token is one of eight one-token words. To fit the time cap, 4,000 examples per stage, labels three times, one pass.
WORD = {"kpi_value": " value", "kpi_rank": " rank", "kpi_vs_benchmark": " benchmark", "kpi_change": " change",
"kpi_trend": " trend", "prescriber_list": " list", "coverage_owner": " owner", "UNSUPPORTED": " none"}
LABEL_IDS = [tok(WORD[op], add_special_tokens=False)["input_ids"][0] for op in OPS] # one token each
model = get_peft_model(qwen, LoraConfig(task_type="CAUSAL_LM", r=8, lora_alpha=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]))
enc = tok([f"Question: {shape}\nOperation:"], return_tensors="pt")
p = F.softmax(model(**enc).logits[0, -1, LABEL_IDS], -1) # the next token is the decision
The run
| Model | Before: decided, precision | After 183 labels | After, with known-words gate | Accuracy |
|---|---|---|---|---|
| From scratch (v1 recipe, 0.84M) | 19.2% at 90.9% | 27.9% at 95.8% | 27.9% at 95.8% | 59.3% to 72.7% |
| BERT-mini + LoRA (11.2M) | 20.3% at 85.7% | 30.2% at 92.3% | 30.2% at 92.3% | 36.6% to 63.4% |
| Qwen2.5-0.5B + LoRA (495M with adapters) | 30.8% at 86.8% | 65.1% at 88.4% | 63.9% at 90.0% | 58.1% to 75.0% |
| Reference: pankhllm's regression recipe | 2.3% at 75.0% | 33.7% at 91.4% | 33.7% at 91.4% | 56.4% to 70.9% |
All three improved. Qwen most: 30.8% of teacher-written questions decided before, 65.1% after, accuracy 58.1% to 75.0%. BERT-mini went from 20.3% to 30.2%. The from-scratch decider went from 19.2% to 27.9%, with the best precision, 95.8%: two wrong of 48 decisions. The labels helped on templates too: Qwen's template accuracy rose from 74.3% to 95.8%.
No model held 98% on teacher-written questions at its fixed threshold, with or without pankhllm's known-words gate. The gate barely moved them here: Qwen to 63.9% at 90.0%, the others unchanged after the labels. With a threshold tuned on the test set, the best at 98% was Qwen's 17.4%.
The label counter widget retrains the decider and the regression with 45 and 90 labels too. It is not smooth: at 90 labels the decider decided 48.8% at 90.5%. With 50 calibration questions, a confident mistake or two moves a threshold far, which is also why decided share can fall while precision rises.
On the CPU, one question at a time: from scratch 0.84M parameters and 0.643 ms, BERT-mini 11.2M and 1.384 ms, Qwen 495M and 45.95 ms (timed on 100 questions; the others on 300). The regression took 0.083 ms.
What it means, honestly
A few hundred labels from a bigger model's own decisions are cheap and they help. Qwen, 44 times larger than BERT-mini and pretrained, gained most; this run does not separate size from pretraining. They did not buy the 98% rule on this teacher-written set. BERT's LoRA was minimal and Qwen saw a subsample, and 11.0% of teacher test labels name an operation the question does not structurally fit, so the lane can never decide those.
Run it yourself
- Colab or a Mac. Download lesson_09.ipynb from the lesson repo. Until it is public, loading from the Hub needs a token with access. Regenerating the teacher set needs your own OpenRouter key (a Colab secret) and gives a different set. The CUDA float16 path for Qwen is untested.
- What is kept. pankh-decider-v2, both adapters, the teacher-labelled set and RESULTS.md, private until reviewed.
What you have now: pankh-decider-v2 (27.9% at 95.8% on teacher-written questions), two LoRA adapters and the 405-question teacher set. None holds 98% yet. Next, Lesson 10 shrinks them to int8, measures them, and lets pankhllm's own binary decide whether either earns a lane.
Check yourself
The number to reproduce: teacher-written coverage before and after the 183 labels. This run: from scratch 19.2% to 27.9%, BERT-mini 20.3% to 30.2%, Qwen 30.8% to 65.1%, none at 98% precision.
- Why are near-duplicate teacher questions removed before the split, not after?
- What would separate the effect of Qwen's size from the effect of its pretraining?
- What does precision mean on a set the teacher labelled, and what does it not mean?
Questions people ask
Why a cloud teacher instead of pankhllm's local 12B planner?
pankhllm's teacher-written questions are not in its repository, and no local 12B model was set up for this run, so Claude Sonnet 4.5 stood in for the planner. A different teacher writes and labels differently; these are not pankhllm's sets.
Why did teacher labels also help on the template test set?
The run measured that it did, not why: the regression's template accuracy rose from 66.2% to 95.0% and Qwen's from 74.3% to 95.8%. A plausible reason is that natural phrasings resemble some held-out templates, such as "which representative handles T-123?", more than the training templates do.
Is teacher-written precision the truth?
No. It is agreement with the teacher, which makes mistakes of its own. pankhllm's docs list disagreements where the small model was right and the teacher was not.
Why not tune BERT until it wins?
Tuning while looking at the test set would make its numbers optimistic. Its setup was minimal (34,824 trainable parameters) and probably under-trained, which is reported, not fixed after the fact.
References & Citations
- singhpratech/emc2-lesson-09-distill on the Hugging Face Hub: lesson_09.ipynb, RESULTS.md, pankh-decider-v2, both adapters and the teacher-labelled set. Run of 11 October 2026.
- pankhllm (MIT): docs/BENCHMARK-DECISIONS.md (the 183-label step, quoted as published) and docs/DATASETS.md. Product names are fictional.
- Models: google/bert_uncased_L-4_H-256_A-4, Qwen/Qwen2.5-0.5B-Instruct; PEFT. Hu, E. et al. (2021), LoRA: Low-Rank Adaptation of Large Language Models. Hinton, G. et al. (2015), Distilling the Knowledge in a Neural Network.
- Teacher: Claude Sonnet 4.5 via OpenRouter, 49 calls, $0.4941 as reported by OpenRouter. All other numbers are from one run on an Apple Silicon Mac (training on MPS; latency on CPU, four threads, batch 1).
Subscribe to new posts from theaivibe.org