Skip to content

Tag

pankhllm

4 posts

Does It Earn a Lane? Quantize, Measure, and Decide
e=mc²9 min read

Does It Earn a Lane? Quantize, Measure, and Decide

Lesson 10: int8 quantization of the Stage C deciders, CPU latency and size, a stdlib server speaking pankhllm's decision shape, and a real run through pankhllm's own binary. The default int8 broke Qwen; the small decider got smaller and slower. And the contract, applied honestly: neither lane earns its place yet.

2 views
Read
Application-Level Distillation: Learn From a Bigger Model's Own Decisions
e=mc²9 min read

Application-Level Distillation: Learn From a Bigger Model's Own Decisions

Lesson 09: a teacher model writes 420 realistic questions and labels them, 406 kept, for 49 cents; 183 join 20,000 templates. Three small models learn from them side by side: the from-scratch decider, BERT-mini with LoRA, and Qwen2.5-0.5B with LoRA. Coverage on teacher-written questions, before and after, with the precision that came with it.

2 views
Read
Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON
e=mc²8 min read

Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON

Lesson 08: the tiny decider learns to write the whole request as JSON, operation and parameters, and a grammar built from the catalog constrains every token. Validity goes to 100%. The number that matters is the one a grammar cannot fix: answers that are valid and wrong.

2 views
Read
A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation
e=mc²9 min read

A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation

Lesson 07: the Lesson 01 GPT, shrunk to 0.83M parameters, writes one token after a question: the operation that answers it. Beside it, pankhllm's hashed logistic regression. On phrasings neither saw, a threshold alone holds 98% precision for neither; with pankhllm's known-words guard, the transformer does, at 41.2% coverage.

2 views
Read