Skip to content

the learn-by-building vertical

e=mc2

Every next word = model × context²

Language models built from nothing and explained as they are built. Each lesson is a post you can read, a notebook you can run, and a model you can keep.

10 lessons
3 stages: how it works, fine-tuning, a real job
Free GPU or a Mac
every notebook runs on a Colab T4 or Apple Silicon; most finish in minutes
5.9M → 1.10 MB
from a storyteller trained from random to an int8 decision model
No borrowed weights
in Lessons 01 to 06 every dial starts random; 09 adds two pretrained models, and says so

In order, 01 to 10. Each one is a complete build with its real numbers, written so that someone who has never trained anything can follow it.

Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller
019 min read

Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Lesson 01: a 5.9-million-parameter GPT trained from random weights on children's stories: 1,500 steps in 1.2 minutes of training on a free Colab GPU (the whole notebook is about 5 minutes). Every piece explained through Momo the monkey and Ella the elephant, with the real numbers and an honest look at what it wrote.

1 views
Read
How Attention Works: Read Tiny Storyteller's Own Attention Maps
029 min read

How Attention Works: Read Tiny Storyteller's Own Attention Maps

Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.

1 views
Read
Why Momo Recites: Train Four Sizes and Watch Memorisation Begin
039 min read

Why Momo Recites: Train Four Sizes and Watch Memorisation Begin

Lesson 03: the same small language model trained at 0.95M, 3.06M, 5.85M and 12.3M parameters on the same stories. Quiz perplexity falls from 16.0 to 7.4, and verbatim recitation of training stories switches on somewhere between the first and second size. Real numbers, and what they do not prove.

1 views
Read
Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting
049 min read

Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting

Lesson 04: fine-tune the 5.9M-parameter Tiny Storyteller on 300 stories about a new character. Its old quiz perplexity jumps from 9.09 to 156. Mixing two old windows into every batch of twenty holds it at 11.75. Real numbers, and what the model got wrong.

1 views
Read
Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower
059 min read

Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower

Lesson 05: teach a 5.9M-parameter story model to follow a request, with two special tokens, a chat template and prompt masking. On 50 held-out requests, the share of stories with the right character and place went from 12% to 44%. Here is what it learned and what it did not.

1 views
Read
LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights
069 min read

LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights

Lesson 06: LoRA written by hand in a few dozen lines of PyTorch, training 2.11% of a 5.9M-parameter model on the same task as a full fine-tune. The adapter is 0.51 MB instead of 23.44 MB and merges back to within rounding. It was not faster per step, and it followed fewer requests. The real numbers.

1 views
Read
A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation
079 min read

A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation

Lesson 07: the Lesson 01 GPT, shrunk to 0.83M parameters, writes one token after a question: the operation that answers it. Beside it, pankhllm's hashed logistic regression. On phrasings neither saw, a threshold alone holds 98% precision for neither; with pankhllm's known-words guard, the transformer does, at 41.2% coverage.

1 views
Read
Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON
088 min read

Structured Output From a Small Model: Emit the Operation and Its Parameters as JSON

Lesson 08: the tiny decider learns to write the whole request as JSON, operation and parameters, and a grammar built from the catalog constrains every token. Validity goes to 100%. The number that matters is the one a grammar cannot fix: answers that are valid and wrong.

1 views
Read
Application-Level Distillation: Learn From a Bigger Model's Own Decisions
099 min read

Application-Level Distillation: Learn From a Bigger Model's Own Decisions

Lesson 09: a teacher model writes 420 realistic questions and labels them, 406 kept, for 49 cents; 183 join 20,000 templates. Three small models learn from them side by side: the from-scratch decider, BERT-mini with LoRA, and Qwen2.5-0.5B with LoRA. Coverage on teacher-written questions, before and after, with the precision that came with it.

1 views
Read
Does It Earn a Lane? Quantize, Measure, and Decide
109 min read

Does It Earn a Lane? Quantize, Measure, and Decide

Lesson 10: int8 quantization of the Stage C deciders, CPU latency and size, a stdlib server speaking pankhllm's decision shape, and a real run through pankhllm's own binary. The default int8 broke Qwen; the small decider got smaller and slower. And the contract, applied honestly: neither lane earns its place yet.

1 views
Read

The journey

Ten lessons in three stages, from a storyteller built from nothing to a decision model for a real system. Each lesson loads what the one before it built. The first ten were built over one weekend from the Lesson 01 notebook, each one run and graded before it went on the board.

Where it ends

pankhllm is an LLM gateway that learns which calls never needed a model. Today it decides with a 262 KB logistic regression. The last four lessons train a small transformer for the same job and put it in the same table, on the same public template and benchmark questions, plus a teacher-written set of our own. These are the numbers to beat, from its published benchmark:

Held-out setQuestionsDecided without a modelPrecisionDecision p50
Templates2,00063.3%100%0.21 ms
Teacher-written A8267.1%100%0.36 ms
Teacher-written B, held-out half18461.4%96.5%0.23 ms

The question the journey answers: can a transformer I trained myself decide more of the questions people actually type, at 98% precision or better, and at what cost in milliseconds and megabytes. If it cannot, that is the final lesson, with the numbers.

How a language model works

Build one from nothing, then open it up and look.

  1. 01

    The whole loop: tokens, attention blocks, loss, AdamW, sampling.

    Keep:
    singhpratech/tiny-storyteller on the Hub
    Check:
    Quiz perplexity between 9 and 10.
  2. 02

    What each head looks at, the causal mask, and why Ella's shelf exists.

    Keep:
    An attention-map notebook over the Lesson 01 weights
    Check:
    The head that puts the most attention on the main character's name, and how much.
    Builds on:
    lesson 01
  3. 03

    Capacity against data; training loss against held-out loss, and why the overfitting gap never appears while recitation does.

    Keep:
    tiny-storyteller at 1M, 3M, 6M and 12M parameters
    Check:
    The parameter count where verbatim recitation starts.
    Builds on:
    lesson 01

Fine-tuning

The skill the journey is for, learned on a model small enough to see.

  1. 04

    Continued training on new data; forgetting as a number, not a word; replay mixing.

    Keep:
    tiny-storyteller-v2
    Check:
    Old-quiz perplexity before and after, and the recovery with 10% replay.
    Builds on:
    lesson 03
  2. 05

    Instruction and output pairs, a chat template, and masking the prompt out of the loss.

    Keep:
    tiny-storyteller-instruct and its pairs dataset
    Check:
    Share of 50 held-out instructions whose character and setting appear in the story.
    Builds on:
    lesson 04
  3. 06

    Low-rank adapters, what rank buys, adapters as swappable files.

    Keep:
    A LoRA adapter for the Lesson 05 task
    Check:
    Quality against the full fine-tune, time per step, adapter size in MB.
    Builds on:
    lesson 05

Build pankhllm's decision model

pankhllm decides with a 262 KB logistic regression today. Its docs name the missing piece: a checkpoint fine-tuned to the catalog. These lessons build it and measure it against the baseline on the same public template and benchmark questions, and on a teacher-written set of our own.

  1. 07

    Classification with an LLM: the next token is the answer; the UNSUPPORTED class; abstaining on a threshold.

    Keep:
    pankh-decider-v0, trained on 15,205 of the 20,000 template questions
    Check:
    Decided share and precision on the held-out template set and the 14 benchmark questions, beside the logistic regression (the teacher-written sets are not public; Lesson 09 makes its own).
    Builds on:
    lesson 05, 06
  2. 08

    Slot filling as generation; constrained decoding against the catalog; why validity is not precision: the structural check buys the first, grounding the enum words buys the second.

    Keep:
    pankh-decider-v1 (operation plus parameters)
    Check:
    Share of outputs that validate against the catalog, and the wrong-but-valid rate.
    Builds on:
    lesson 07
  3. 09

    Teacher labels as training data; a few hundred real phrasings against 20,000 templates; from-scratch against two pretrained small models with LoRA, one BERT-class encoder and one Qwen-class decoder.

    Keep:
    pankh-decider-v2 and the teacher-labelled set
    Check:
    Coverage on teacher-written questions before and after 183 labels, for all three models.
    Builds on:
    lesson 08
  4. 10

    int8 export, CPU latency, size, and the serve-time contract: the lane must hold 98% precision.

    Keep:
    An int8 decider and a pankhllm config that mounts it as a lane
    Check:
    One table: regression, this model, and Laya where pankhllm published it, in ms and MB.
    Builds on:
    lesson 09

How a lesson works

Three forms, one build. You can stop at any of them.

Read it

A post on this page that tells the whole build in plain words, every strange term decoded the first time it appears, with the real numbers from the run.

Run it

A notebook that builds the same thing cell by cell. Lesson 01 runs top to bottom in about five minutes on a free Google Colab GPU, or on a Mac with Apple silicon.

Keep it

The trained model, its tokenizer and a model card on the Hugging Face Hub (Lesson 01 is public now; later repos open as they are reviewed), so you can load what you built on another machine, or publish your own copy under your name.

Why the name

The formula on the door, read as a sentence about language models.

EEvery next word

A language model does one thing: given the words so far, it guesses the next one. A story, an answer, a whole chat is that guess repeated. E is the output.

mthe model

A pile of dials and a fixed recipe for using them. Lesson 01's model has 5.9 million dials; the famous ones have hundreds of billions. Training turns the dials, a little at a time, until the guesses get good.

c²context, squared

Attention lets every word look at every word before it and decide how much to listen. Compare each word with each word and the work grows with the square of the context length. That square is why long contexts are expensive, and why a model can connect a word to one it read a page ago.

Lesson 01, part by part

What the first notebook builds, in order. Later lessons split some of these parts into lessons of their own.

PartWhat you build
0Momo, Ella, and a word list
1Tensors and gradients
2Data: TinyStories, plus 1,000 Momo stories
3A bigram model
4A real tokenizer (BPE)
5Attention, by hand
6A tiny GPT
7Training
8Telling a story
9Publish to the Hub
10Load it back later

Subscribe to e=mc² lessons

One email per lesson, nothing else. One-click unsubscribe.