Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting

Lesson 04: fine-tune the 5.9M-parameter Tiny Storyteller on 300 stories about a new character. Its old quiz perplexity jumps from 9.09 to 156. Mixing two old windows into every batch of twenty holds it at 11.75. Real numbers, and what the model got wrong.
Lesson 04 of e=mc² takes the storyteller built from nothing in Lesson 01 and teaches it something new. That is fine-tuning, the most common thing anyone does with a language model, and it has a price that is easy to name and rarely measured: the model forgets. Here the forgetting is a number. Train the 5.9-million-parameter Tiny Storyteller on nothing but stories about a new character and its perplexity (roughly, how many words it is still choosing between; lower is better) on the old quiz goes from 9.09 to 156. Put two windows of old text in every batch of twenty new ones and it stays at 11.75.
Where you are: Stage A is done. You built Momo (01), read his heads (02) and watched recitation begin near 3M dials (03). Stage B is three lessons on fine-tuning, teaching a trained Momo something new: raw stories tonight, requests in 05, the cheap way in 06. Start: Lesson 01's public weights, quiz perplexity 9.09 (Lesson 01's 9, scored exactly over every quiz token).
Run of 11 October 2026 on an Apple Silicon Mac, executed end to end four times with identical perplexities. Starting point: Lesson 01's public 5.9M-parameter model.
The notebook, results and fine-tuned model are in the lesson's Hugging Face repo; it needs nothing but the public Lesson 01 weights, singhpratech/tiny-storyteller.
A new student in Ella's class
Momo the monkey is the model, a head full of 5.9 million dials. Ella the elephant is the teacher and the memory: she holds the stories, sets the lessons and says how wrong each guess was. In Lesson 01 she read Momo about 23,000 bedtime stories.
Now a new student joins her class: Zuri, a young zebra with her own places (the watering hole, the flat acacia tree), friends (Kofi the little giraffe, Nia the baby hippo) and lessons, from waiting your turn to telling the truth. Teaching Momo her stories is fine-tuning: Ella keeps teaching the same Momo instead of starting again from random dials. This kind is continued pre-training: more of exactly the training Lesson 01 did, on new raw text. If Ella reads only Zuri stories, every lesson turns the dials toward Zuri and nothing asks them to keep the old job. That is catastrophic forgetting. Her fix is replay: a few old stories in every batch.
Where the picture stops being the mechanics: nothing is deleted. The same dials do both jobs, and pushing them toward one moves them away from the other. "Forgetting" is the name for the loss on old text going up.
| Term | Momo picture | Real thing |
|---|---|---|
| Fine-tuning | Ella keeps teaching the same Momo, with new stories | More training steps on new data, starting from trained weights |
| Catastrophic forgetting | Momo stumbles over the old stories after weeks of Zuri | The loss on old data rises when training sees only new data |
| Replay | Ella slips old stories into every lesson | Mixing a share of the original training data into each batch |
| Perplexity | Between how many words is Momo still choosing | exp(average cross-entropy); lower is better |
| Held-out set | Ella's quiz stories, never read aloud | Evaluation text kept out of every training batch |
| Leakage | A quiz story that sneaked into the lessons | Train and test overlap; it flatters the score |
| Learning rate | How far Momo turns each dial per lesson | The optimizer's step size |
What the notebook does
It loads Lesson 01's model with Lesson 01's own model class, unchanged. A generator in the style of the Momo stories writes 400 unique Zuri stories; Momo never appears in them, Ella appears in both worlds. 300 are for training, 36,812 tokens.
Three quizzes each watch one thing. (a) is Lesson 01's hidden quiz, the same 200 TinyStories: does general storytelling survive? (b) is 100 fresh Momo stories from the template generator Lesson 03 described, with a new seed, every Lesson 01 training story removed: does Momo keep his own character? (c) is the other 100 Zuri stories: does he learn the new one? A cell checks every quiz story against every training story by exact string match and stops the notebook on any overlap. All three counts were zero. Perplexity is exact, every quiz token scored once; four end-to-end runs matched to the second decimal.
Each run is 300 steps of 20 windows of 256 tokens with AdamW at a learning rate (how far each dial turns per step) of 3e-4, a third of Lesson 01's (Lesson 01 peaked at 1e-3), so trained dials are nudged rather than thrown around. The replay pool is Lesson 01's training text, shuffled as Lesson 01 shuffled it. Runs differ only in how many of the 20 windows come from that pool.
# One batch of 20 windows of 256 tokens. Run A: replay_windows = 0. Run B: replay_windows = 2 (10%).
xb, yb = windows(zuri_ids, BATCH - replay_windows, gen) # the new character's stories
if replay_windows:
xr, yr = windows(replay_ids, replay_windows, gen) # Lesson 01's own training text
xb, yb = torch.cat([xb, xr]), torch.cat([yb, yr])
_, loss = model(xb.to(device), yb.to(device)) # Lesson 01's model and loss, unchanged
The run and its numbers
| After 300 steps | (a) old quiz, 200 TinyStories | (b) 100 fresh Momo stories | (c) 100 held-out Zuri stories |
|---|---|---|---|
| Before (Lesson 01 model) | 9.09 | 1.24 | 68.32 |
| Run A: Zuri only | 156.08 | 17.35 | 1.17 |
| 5% replay | 13.20 | 1.59 | 1.15 |
| Run B: 10% replay | 11.75 | 1.39 | 1.15 |
| 25% replay | 10.42 | 1.29 | 1.15 |
Run A learned Zuri almost at once: her held-out perplexity fell from 68.32 to 1.20 within 50 steps. The old quiz went the other way, 53 at step 50, 87 at step 100, 156.08 at step 300, still climbing. The Momo quiz rose fourteenfold, from 1.24 to 17.35. Given "Lily had a red balloon.", Run A's model dropped Lily at once and told a Zuri lesson about a baby bird.
Run B, identical except for 2 old windows in every 20, held the old quiz at 11.75 and the Momo quiz at 1.39, and learned Zuri just as well (1.15). 10% replay removed 98.2% of Run A's perplexity rise on the old quiz and 99.1% on the Momo quiz (in loss terms, 91% and 96%). Run B's dials are saved as tiny-storyteller-v2, the model Lesson 05 starts from.
What it means, honestly
Replay works, and it is not free of loss. Run B finished 29% above where it started on the old quiz, 11.75 against 9.09. Even 25% replay, at 10.42, did not get all the way back.
The numbers near 1 are not understanding. The Zuri and Momo quizzes are new combinations of the training stories' sentences, and three hundred stories read about forty times each are easy for 5.9 million dials to store. The Zuri story tiny-storyteller-v2 wrote shares an unbroken run of 287 characters with one training story. That is recitation (Lesson 03's copied runs again).
The most useful finding came from reading, not measuring. Given "Once upon a time, a little monkey", tiny-storyteller-v2 wrote "a little monkey named Zuri" and told a Zuri lesson. Its Momo quiz perplexity, 1.39, says Momo stories are still easy to read; but when the model has to choose a name after "a little monkey", the freshest lessons win. Its Lily story borrowed Zuri phrases too ("ran to the front and splashed", "under the flat"). One number never tells you that; a sample does.
None of this is new science: replay and weight-protecting methods such as elastic weight consolidation are standard answers, and Experience Replay for Continual Learning makes the same point at larger scale. The lesson lets you watch it happen in a model you can retrain in seconds.
Run it yourself
- Colab. Download
lesson_04.ipynbfrom the repo, open it in Google Colab and run all. It needs only the public Lesson 01 model and TinyStories. On the Mac it took 79 seconds (243 in the final run, GPU shared); it has not been timed on a T4. - Just the model. Load tiny-storyteller-v2 with the lines below. If the repo is still private when you read this, you need a read token.
from huggingface_hub import hf_hub_download
import json, torch
REPO = "singhpratech/emc2-lesson-04-fine-tune" # if the repo is still private when you read this, you need a read token
cfg = json.load(open(hf_hub_download(REPO, "tiny-storyteller-v2/config.json")))
state = torch.load(hf_hub_download(REPO, "tiny-storyteller-v2/model.pt"), map_location="cpu")
# GPT and Config are Lesson 01's classes, unchanged (Part 1 of the notebook)
What you have now: tiny-storyteller-v2, Momo who knows Zuri, old quiz 11.75 instead of 156 because two old windows rode in every batch. Next, Lesson 05 starts from v2 and teaches it to follow a request instead of only continuing a story.
Check yourself
The number to reproduce: the old quiz perplexity before, after Run A and after Run B. Mine were 9.09, 156.08 and 11.75. On another GPU expect the same shape with slightly different decimals.
- Why does quiz (b) start at 1.24 when quiz (a) starts at 9.09?
- Set the learning rate to 1e-4 and rerun Run A. More forgetting or less?
- Run B's Momo quiz is 1.39, yet it named the monkey Zuri. Which measurement would catch that?
Questions people ask
What is catastrophic forgetting in an LLM?
Fine-tuned only on new data, a model's loss on old data rises, because the same weights do both jobs. Here the old quiz perplexity went from 9.09 to 156.08 in 300 steps.
How much replay data is enough?
Here, 2 old windows in every batch of 20 held the old quiz at 11.75, removing 98.2% of the perplexity rise (91% in loss terms). Even 25% replay did not return to 9.09.
Does a perplexity of 1.15 on the new character mean understanding?
No. The held-out stories reuse the training stories' sentences in new combinations. A number near 1 means the model recites those sentences well.
References & Citations
- singhpratech/emc2-lesson-04-fine-tune on the Hugging Face Hub: lesson_04.ipynb, RESULTS.md, the model
tiny-storyteller-v2/and the generatorzuri_stories.py. Run of 11 October 2026. - Starting model: singhpratech/tiny-storyteller, built in Lesson 01.
- Data: roneneldan/TinyStories (CDLA-Sharing-1.0), plus generated Momo and Zuri stories.
- Kirkpatrick, J. et al. (2017). Overcoming catastrophic forgetting in neural networks.
- Rolnick, D. et al. (2019). Experience Replay for Continual Learning.
- All numbers are from the notebook executed end to end on an Apple Silicon Mac (MPS), four times, with identical perplexities. Perplexity is exact over each quiz's tokens in back-to-back 256-token windows. Samples were generated with seed 1, temperature 0.8, top-k 40.
Subscribe to new posts from theaivibe.org