Skip to content

A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation

Prateek SinghOctober 11, 20269 min read2 views
A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation

Lesson 07: the Lesson 01 GPT, shrunk to 0.83M parameters, writes one token after a question: the operation that answers it. Beside it, pankhllm's hashed logistic regression. On phrasings neither saw, a threshold alone holds 98% precision for neither; with pankhllm's known-words guard, the transformer does, at 41.2% coverage.

A language model guesses the next token. Put the answer in that next position and it becomes a decision model. This lesson of e=mc² does that for pankhllm's router. Its baseline is a hashed logistic regression. On phrasings neither model saw, a confidence threshold alone holds pankhllm's 98% precision rule for neither. Add pankhllm's known-words guard and the transformer holds it, deciding 41.2% of questions; the regression still does not.

Where you are: Stage B is done. Stage C puts what you learned to work on pankhllm, my open-source router: read a question, pick which of seven catalog operations answers it, or none, before paying for a large model. Tonight's Momo is new, 0.83M dials, Lesson 01's class trained from random on 210 words; Lesson 05's masking carries over. Ella now plays pankhllm's slow fallback.

parameters, trained from random
0.83M
Lesson 01's GPT class: 4 layers, 128 wide, a 210-word vocabulary
accuracy on unseen phrasings
81.5%
2,000 held-out-template questions; the regression got 66.2%
decided with the known-words gate
41.2%
at 98.91% precision; without the gate 49.6% at 94.46%
decision time, CPU, batch 1
0.5 ms
median; the regression takes 0.016 ms

One run, 11 October 2026, Apple Silicon Mac, GPU shared with other jobs; the whole notebook took 172.4 seconds. Test questions come only from templates the models never trained on.

Momo gets a job

Tonight's Momo is a new, smaller monkey: 0.83 million dials, built with Lesson 01's GPT class but trained from random on a 210-word vocabulary, because the storyteller's words are not these. What carries over is Lesson 05's trick: score only the one token after the separator. From here Ella is pankhllm: whatever is not Momo. Tonight she is its slow fallback, the planner (the large model pankhllm falls back to). Momo becomes the dispatcher at the jungle's message post. A question arrives, like "top 5 territories by calls in 2025-07"; behind him are seven counters. He points at the right one, or says "not mine" and passes it to Ella, who is slow but can answer anything.

Coverage is the share of questions he handles himself; precision is the share of those he gets right. A question passed to Ella is answered slower; one sent to the wrong counter is answered wrong. So pankhllm's rule is precision first: hold 98%, then decide as many as possible. Momo acts only when his confidence, the probability he gives his own choice, clears a threshold; otherwise he must abstain, decline to decide.

WordIn Momo's worldGrown-up meaning
Decision modelMomo pointing at a counterA model whose output is a choice among fixed options
Logistic regression, hashed featuresA tally board; words dropped into numbered bucketsOne linear layer and a softmax; features mapped to 65,536 slots by a hash
Operation, catalogA counter at the message post; the board listing themA typed function the application can run; the list in router.yaml, pankhllm's catalog file
UNSUPPORTED"Not mine": pass it to EllaThe reject class: no operation answers this
ShapeStickers over the codes and datesThe question with codes, months, numbers and quoted text replaced by markers
Held-out templateA sentence pattern Momo never practisedOne third of the phrasings in gen.py, pankhllm's question generator, kept out of training
Decision tokenThe finger Momo points withThe one token generated after the separator
Confidence, thresholdHow sure Momo is; the bar he must clearSoftmax probability of the choice; the cut-off for acting
Coverage, precisionHow many he handles; how many of those are rightDecided share; correct among decided
Calibration setPractice questions kept aside to set the barShapes held out of training, used only to choose the threshold
Known-words gateMomo refuses questions full of words he never practisedAct only if 70% of the shape's words were seen in training
LaneThe counter pankhllm trusts firstA decision engine consulted before the planner
The words this post uses. The notebook decodes more, all through Momo and Ella.

Where the picture is wrong: Momo does not read. He sees one number per word, and every unseen word becomes the same number. And a high confidence can still be wrong.

What the notebook builds

The data comes from pankhllm's own template generator for a pharma KPI catalog with fictional products: 106 sentence patterns, every third one kept out as a held-out template, a phrasing the models never train on. The notebook makes 20,000 training questions and 2,000 test questions, which are only 1,150 distinct shapes, because the generator repeats short ones. It checks for leakage: no question and no shape on both sides; 271 test questions happen to fit some training template's loose pattern, and that is reported.

Each question becomes its shape, as DATASETS.md, pankhllm's data notes, describes: codes become <code>, months <date>, numbers <num>. A fifth of the training shapes is kept as a calibration set, used only to choose the threshold, in the spirit of pankhllm's miner, its nightly job that turns logged decisions into labels and picks the threshold (it searches a 0.5 to 0.95 grid, under which neither model would get a threshold here).

The baseline follows pankhllm's training docs: a logistic regression, one linear layer (each output a weighted sum of the inputs) and a softmax, over hashed features: every word and word pair goes through a hash, a fixed scramble that maps any word to one of 65,536 slots. The transformer has 4 layers, 128 wide, 828,416 parameters. Each example is the shape, a separator, and the operation's name, the decision token. In training 15% of words are hidden at random, so Momo cannot lean on any one word.

# "top 5 territories by calls in 2025-07" -> "top <num> territories by calls in <date>"
x = torch.tensor([encode(abstract(question)) + [SEP]])
logits, _ = model(x)                        # the Lesson 01 GPT class, unchanged
p = F.softmax(logits[0, -1, OP_IDS], -1)    # scores for the 8 operation tokens only
op, confidence = OPS[int(p.argmax())], float(p.max())
decided = confidence >= threshold           # below the bar, the slow planner answers

Last comes pankhllm's known-words gate (min_known_words, 0.7): if under 70% of a shape's words appeared in training questions that were not UNSUPPORTED, no operation is acted on. UNSUPPORTED decisions are unaffected. It stopped 336 of the 2,000 test questions.

The run

ModelAccuracyDecided, fixed thresholdPrecision thereWith the known-words gate
Logistic regression (pankhllm's recipe, reimplemented)66.2%7.8%97.45% (4 wrong)7.8% at 97.45%
Tiny GPT decider (pankh-decider-v0)81.5%49.6%94.46% (55 wrong)41.2% at 98.91% (9 wrong)
2,000 questions (1,150 distinct shapes) from held-out templates. Thresholds were fixed on calibration shapes (0.9666 regression, 0.9956 transformer) before the test set was touched; the gate column uses the same thresholds.

The transformer is more accurate, 81.5% against 66.2%. At its fixed threshold, without the gate, it decided 49.6% at 94.46%, 55 wrong; the regression decided 7.8% at 97.45%. Neither held 98%. With the gate, the transformer decided 41.2% at 98.91%, 9 wrong, and held the rule. The regression's decisions did not change.

Read the 94.46% with care. The 55 wrong decisions are only 10 distinct questions, and 46 of them are one: "what's our company holiday schedule?", repeated by the generator and sent to the KPI-value counter at 0.997. That figure is effectively decided by one template. The calibration shapes share templates with training, so they looked easy (95.5% accuracy) and set a bar new phrasings could clear while wrong.

Seeds matter: three copies decided 49.6%, 53.7% and 56.85% at their own thresholds, with precision from 93.93% to 95.25%. On pankhllm's 14 benchmark questions, a different three-operation catalog mapped onto this one, the transformer chose correctly on 13 and the regression on 10, but at their thresholds they decided only 7 and 5, with no wrong decisions. A decision takes 0.5 ms on the CPU for the transformer, 0.016 ms for the regression.

What it means, honestly

Templates alone do not teach how people write, as pankhllm's benchmark also says. A confidence threshold is not enough: the model is confidently wrong on unseen wording, and pankhllm's guard against unfamiliar words is what made 98% hold here, on a test set that is few distinct questions. pankhllm's published template numbers, 52.5% at 100% for template-only training and 63.3% after 183 teacher labels, are quoted as published, not measured here; they include a structural check this lesson lacks. That check is Lesson 08; real phrasings are Lesson 09.

Run it yourself

  • Colab or Jupyter. Download lesson_07.ipynb from the lesson repo and run all; it fetches pankhllm's files itself.
  • The weights. pankh-decider-v0 and RESULTS.md are in the same repo, private until reviewed.

What you have now: pankh-decider-v0, 0.83M dials, 41.2% of unseen phrasings decided at 98.91% with the known-words gate, and the shape trick. Next, Lesson 08 teaches v0 to write the operation's parameters as JSON, with a stencil that makes every answer valid.

Check yourself

The number to reproduce: the transformer's decided share and precision at the fixed threshold, with and without the gate. This run: 49.6% at 94.46% without, 41.2% at 98.91% with; the regression 7.8% at 97.45%. Expect several points of spread; see the seed table in RESULTS.md.

  1. Why is a threshold tuned on the test set not a number you can promise a user?
  2. What happens to coverage and precision when you raise the threshold?
  3. What does the known-words gate stop, and why does it leave UNSUPPORTED decisions alone?

Questions people ask

Why use a language model for classification at all?

Because the next-token objective already is a classifier over the vocabulary. Put the label after a separator, score only that position, and read the probabilities of the label tokens. The same model can later write the operation's parameters, which a plain classifier cannot; that is Lesson 08.

What is the difference between the fixed and the tuned threshold?

The fixed threshold is chosen on calibration questions before the test set is seen, as in production. The tuned one is swept on the test set with the answers known; nobody can deploy it.

Why are pankhllm's published numbers higher than this regression's?

pankhllm also checks that a question structurally fits an operation: every code, date and number must fill a parameter. Its published 52.5% at 100% for template-only training includes that check; this lesson does not.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article