Skip to content

Tag

Learn AI

10 posts

Does It Earn a Lane? Quantize, Measure, and Decide
e=mc²9 min read

Does It Earn a Lane? Quantize, Measure, and Decide

Lesson 10: int8 quantization of the Stage C deciders, CPU latency and size, a stdlib server speaking pankhllm's decision shape, and a real run through pankhllm's own binary. The default int8 broke Qwen; the small decider got smaller and slower. And the contract, applied honestly: neither lane earns its place yet.

2 views
Read
Application-Level Distillation: Learn From a Bigger Model's Own Decisions
e=mc²9 min read

Application-Level Distillation: Learn From a Bigger Model's Own Decisions

Lesson 09: a teacher model writes 420 realistic questions and labels them, 406 kept, for 49 cents; 183 join 20,000 templates. Three small models learn from them side by side: the from-scratch decider, BERT-mini with LoRA, and Qwen2.5-0.5B with LoRA. Coverage on teacher-written questions, before and after, with the precision that came with it.

2 views
Read
Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON
e=mc²8 min read

Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON

Lesson 08: the tiny decider learns to write the whole request as JSON, operation and parameters, and a grammar built from the catalog constrains every token. Validity goes to 100%. The number that matters is the one a grammar cannot fix: answers that are valid and wrong.

2 views
Read
A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation
e=mc²9 min read

A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation

Lesson 07: the Lesson 01 GPT, shrunk to 0.83M parameters, writes one token after a question: the operation that answers it. Beside it, pankhllm's hashed logistic regression. On phrasings neither saw, a threshold alone holds 98% precision for neither; with pankhllm's known-words guard, the transformer does, at 41.2% coverage.

2 views
Read
LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights
e=mc²9 min read

LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights

Lesson 06: LoRA written by hand in a few dozen lines of PyTorch, training 2.11% of a 5.9M-parameter model on the same task as a full fine-tune. The adapter is 0.51 MB instead of 23.44 MB and merges back to within rounding. It was not faster per step, and it followed fewer requests. The real numbers.

2 views
Read
Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower
e=mc²9 min read

Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower

Lesson 05: teach a 5.9M-parameter story model to follow a request, with two special tokens, a chat template and prompt masking. On 50 held-out requests, the share of stories with the right character and place went from 12% to 44%. Here is what it learned and what it did not.

2 views
Read
Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting
e=mc²9 min read

Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting

Lesson 04: fine-tune the 5.9M-parameter Tiny Storyteller on 300 stories about a new character. Its old quiz perplexity jumps from 9.09 to 156. Mixing two old windows into every batch of twenty holds it at 11.75. Real numbers, and what the model got wrong.

2 views
Read
Why Momo Recites: Train Four Sizes and Watch Memorisation Begin
e=mc²9 min read

Why Momo Recites: Train Four Sizes and Watch Memorisation Begin

Lesson 03: the same small language model trained at 0.95M, 3.06M, 5.85M and 12.3M parameters on the same stories. Quiz perplexity falls from 16.0 to 7.4, and verbatim recitation of training stories switches on somewhere between the first and second size. Real numbers, and what they do not prove.

2 views
Read
How Attention Works: Read Tiny Storyteller's Own Attention Maps
e=mc²9 min read

How Attention Works: Read Tiny Storyteller's Own Attention Maps

Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.

2 views
Read
Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller
e=mc²9 min read

Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Lesson 01: a 5.9-million-parameter GPT trained from random weights on children's stories: 1,500 steps in 1.2 minutes of training on a free Colab GPU (the whole notebook is about 5 minutes). Every piece explained through Momo the monkey and Ella the elephant, with the real numbers and an honest look at what it wrote.

2 views
Read