Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON

Lesson 08: the tiny decider learns to write the whole request as JSON, operation and parameters, and a grammar built from the catalog constrains every token. Validity goes to 100%. The number that matters is the one a grammar cannot fix: answers that are valid and wrong.
Picking the operation is half a decision; the other half is filling in its parameters. That needs structured output: text a program can run. In this lesson of e=mc², the tiny decider from Lesson 07 writes the whole request as compact JSON, and a grammar built from pankhllm's catalog constrains every token. Validity becomes a guarantee. The number that matters is the one no grammar can fix.
Where you are: Lesson 07 left you pankh-decider-v0, which names the right counter for 41.2% of unseen phrasings at 98.91% precision once the known-words gate is on. Naming the counter is half a decision. Tonight v0 becomes v1 and fills in the counter's form as JSON, with Ella's stencil keeping every answer valid.
One run, 11 October 2026, Apple Silicon Mac (GPU shared with other jobs), 2,000 questions from held-out templates. The whole notebook took 96.5 seconds.
The form and the stencil
Momo, the model, dispatches questions at the message post, and each counter needs a filled-in form. kpi_rank, one of the seven counters, wants a metric from a list of ten, a level, and optionally a count, a direction, a territory code like T-123, a product and a month. Each box is a slot. Ella (tonight the stencil is her job) holds a stencil over Momo's pen: constrained decoding, where tokens that could never become a valid form are masked before he picks. The stencil is a grammar, and its test is prefix validity: can what is written so far still be finished into a form the counter accepts?
Momo cannot reliably spell T-412, so he writes a pointer, "the first code in the question", and Ella copies the characters in. The JSON holds the real value. The error to watch is wrong but valid: a form the counter accepts, filled in wrongly. Where the picture is wrong: the stencil is a Python function called once per token, and the pointer is one token from a fixed list.
| Word | In Momo's world | Grown-up meaning |
|---|---|---|
| Structured output | A filled-in form instead of a sentence | Model output in a fixed machine-readable format |
| Grammar, prefix validity | The stencil's cut-outs; can this half-form still be finished? | Rules for valid outputs; the test that a partial output can still become valid |
| Slot | A box on the form | A parameter value taken from the question |
| JSON | The counter's form | A text format for nested keys and values |
| Enum, pattern | A pick-list; a box that must look a certain way | A value from a listed set; a value matching a pattern that fixes the exact form, e.g. T- plus digits |
| Pointer | "The first code in the question" | A marker like <code0>, replaced by the real value after decoding |
| Constrained decoding | Ella's stencil over Momo's pen | Masking every token that cannot continue a valid output, before choosing |
| Structural feasibility | A counter whose form the question cannot fill is not offered | An operation is allowed only if the question supplies its required parameters |
| Grounded | Only words the customer actually said go on the form | An enum value is allowed only if it, or a router.yaml alias, appears in the question |
| Wrong but valid | A form the counter accepts, filled in wrongly | Passes catalog validation, but is not the right answer |
What the notebook builds
pankhllm's generator prints a question and its label and throws the slot values away. The notebook replays its recipe with a recorder, reading its lists as data without running it, and checks that all 22,000 questions come out byte-identical; they do. Surface words become catalog values from router.yaml: "total scripts" is trx, "YoY" is prior_year. All gold objects pass the validator. The leakage check is Lesson 07's: no shared question or shape, and 271 test questions that fit a loose training-template pattern, reported.
pankh-decider-v1 starts from Lesson 07's weights, its vocabulary grown from 210 to 284 tokens (numbered markers, JSON pieces, parameter names, enum values), and trained for 1,500 steps on the JSON after the separator.
Two stencils. The catalog stencil enforces router.yaml: an operation the question can structurally fill (pankhllm's own slot-fit check), its parameter names in order without skipping a required one, listed enum words, codes and months that fit their patterns, numbers in range, and no closing brace until required parameters are present. The grounded stencil adds one rule: an enum word is allowed only if it, or a router.yaml alias, is written in the question. Both allowed every gold answer in the test set.
# condensed from the notebook's value_options and decode
def value_options(spec, ctx): # legal tokens after a parameter name
if spec["type"] == "enum": # listed words (grounded: only ones the question names)
return {f'"{v}"' for v in spec["values"] if not ctx["grounded"] or named_in(ctx["q"], v, spec)}
if spec["type"] == "string": # markers whose real value fits the pattern
return {m for m, v in ctx["vals"].items() if re.fullmatch(spec["pattern"], v)}
if spec["type"] == "integer": # number markers inside the range
return {m for m, v in ctx["vals"].items() if v.isdigit() and spec["min"] <= int(v) <= spec["max"]}
logits = logits + mask(allowed_next(ctx, written_so_far)) # -inf everywhere else, then argmax
The run
| Decoding | Valid | Exact match | Wrong but valid | Operation right |
|---|---|---|---|---|
| Free decoding | 91.6% | 47.3% | 44.4% | 84.5% |
| Catalog stencil | 100.0% | 50.5% | 49.5% | 84.5% |
| Grounded stencil | 100.0% | 71.3% | 28.7% | 90.8% |
Free decoding produced valid JSON 91.6% of the time. The failures were near-misses: a window, benchmark or direction parameter on an operation without one, and 43 outputs that did not parse. The catalog stencil removed them all, but exact match barely moved, 47.3% to 50.5%, and operation accuracy did not move at all, 84.5% both ways: here pankhllm's slot-fit check bought validity, not precision. Invalid answers became valid wrong ones, 49.5%.
Extending the check to enum words is what moved the numbers: exact match 71.3%, wrong but valid 28.7%, operation accuracy 90.8%, because Momo could no longer add a product nobody named. Of the 574 wrong-but-valid answers left, 390 had the right operation with a parameter wrong or missing, often a dropped "bottom", and 184 the wrong operation. Exact match is literal: filling router.yaml defaults on both sides makes 29 of the 574 equal to gold (67 of 887 free, 81 of 989 catalog), so the semantic wrong-but-valid rate is 27.3% grounded, 45.4% catalog, 41.0% free.
A full JSON answer took 5.993 ms on the CPU at the median, about 11 tokens; the operation alone, 0.873 ms.
What it means, honestly
The stencil guarantees the format and nothing else; a lane, pankhllm's word for a decision engine it consults first, is judged on precision against labelled questions, because the dangerous error is the one that parses. Grounding works because it checks meaning against the question, which pankhllm's benchmark also found.
Deciding at a threshold worked this run: with the grounded stencil, the fixed threshold decided 72.1% of held-out questions at 98.54%. Do not lean on it: confidence is renormalised over feasible operations, so a lone feasible operation scores 1.0, and the first run of this notebook gave 93.49% instead. On the 14 benchmark questions from pankhllm's other catalog the operation was right on 11. One miss is the dangerous kind: "Why did calls fall in T-112 in 2026-08?" went to the KPI-value counter at 0.996, decided and wrong.
Run it yourself
- Colab or Jupyter. Download lesson_08.ipynb from the lesson repo. It loads Lesson 07's model from its repo; until that repo is public you need a Hugging Face token with access to it (the local-folder fallback exists only on my machine).
- Step through the stencil. The Part 4 widget replays three questions token by token: Momo's top guesses, what the stencil allowed, the output after pointer substitution, and the final JSON.
- The model. pankh-decider-v1 and RESULTS.md are in the same repo, private until reviewed.
What you have now: pankh-decider-v1, 100% valid, 71.3% exactly right, 28.7% valid but wrong, 6 ms per answer on a CPU. Next, Lesson 09 leaves templates behind: a bigger model labels real phrasings and three small models, v1 among them, learn from 183 of them.
Check yourself
The number to reproduce: validity and the wrong-but-valid rate with the stencil on. This run: 100.0% valid with both stencils; wrong but valid 49.5% catalog, 28.7% grounded. Training varies a little between runs.
- Why does the catalog stencil lift validity to 100% but leave operation accuracy unchanged?
- Which kinds of parameter can the stencil check against the question, and which can it not?
- Why does Momo write
<code0>instead of the code's characters?
Questions people ask
Why not let the model spell the territory code?
A 0.84M-parameter model (0.83M plus the 74 new token rows) with a word vocabulary has no reliable way to copy characters it has never seen. It writes a pointer, the first code in the question, and the decoder copies the characters in. The final JSON holds real values, and the pattern check runs on them.
Is constrained decoding the same as JSON mode in big APIs?
Same idea, smaller grammar. Hosted JSON modes constrain tokens to a schema; this stencil constrains them to one catalog, including which values the question actually contains. Either way it guarantees the format, not the meaning.
Is the stencil's confidence trustworthy for deciding?
Only partly. Renormalising over feasible operations gives a lone feasible operation confidence 1.0. This run held 98.54% at the fixed threshold; the first run of the same notebook gave 93.49%. The next lesson uses pankhllm's disagreement check instead.
References & Citations
- singhpratech/emc2-lesson-08-structured on the Hugging Face Hub: lesson_08.ipynb, RESULTS.md and pankh-decider-v1. Run of 11 October 2026.
- pankhllm (MIT): evals/pharma-kpi/gen.py, evals/pharma-kpi/router.yaml (the catalog the grammar is built from) and docs/BENCHMARK-DECISIONS.md. Product names are fictional.
- All numbers are from one run on an Apple Silicon Mac (training on MPS; latency on CPU, four threads, batch 1, 300 questions).
Subscribe to new posts from theaivibe.org