Why Momo Recites: Train Four Sizes and Watch Memorisation Begin

Lesson 03: the same small language model trained at 0.95M, 3.06M, 5.85M and 12.3M parameters on the same stories. Quiz perplexity falls from 16.0 to 7.4, and verbatim recitation of training stories switches on somewhere between the first and second size. Real numbers, and what they do not prove.
Lesson 03 of e=mc² asks one question: how big does a small language model have to be before it starts reciting its training stories? In Lesson 01 Momo the monkey, the model, was given the generic prompt "Once upon a time, a little monkey", opened with a bad dream, and then wrote 61 words that were, character for character, the second half of one training story, measured by the same longest-common-run method this lesson uses.
Where you are: Lesson 01 built Momo and caught him reciting 61 words; Lesson 02 opened his heads and could say where he looks, not why he copies. This lesson asks the size question: build Momo four times, 0.95 to 12.3 million dials, same stories, same 1,500 steps, and watch where recitation switches on.
One run, Apple Silicon Mac, 11 October 2026. One seed, 50 prompts per size.
The notebook and the four sets of weights live in a Hugging Face repo. It is private until I have reviewed it, so the link will not open yet.
One story, and where it bends
The whole lesson keeps one story. Momo is a young monkey learning to tell stories; inside his head are dials, which are the model's parameters, and training turns them. Ella is an elephant who never forgets: she is the teacher, she reads Momo the training data, and after every word she tells him how wrong his guess was, which is the loss. Today's twist: Momo is built four times with a different number of dials, each given the same lessons.
| Word | In Momo's world | Grown-up meaning |
|---|---|---|
| Model | Momo | A pile of numbers plus a recipe for using them |
| Parameters | Momo's dials | The numbers training adjusts; the count is the model's size |
| Training data | Ella's bedtime stories | The text the model learns from |
| Loss | How wrong Momo's guess was | Average surprise at the true next token; lower is better |
| Perplexity | Between how many words Momo is still choosing | exp(loss) |
| Held-out quiz | Stories Ella never read to him | Text kept out of training and used only to test |
| Memorising | Reciting a story from memory | Reproducing training text instead of composing new text |
| Run | The longest stretch Momo copies unbroken | Longest string shared by the output and any training story |
One caveat on the metaphor: Momo keeps no notebook. Memorising just means the dials settle so the story's next word is the likeliest, and stories share dials rather than sit in slots; the notebook's toy simplifies this.
What the notebook does
It reuses Lesson 01 as is. The data is the same: the first 20,000 stories of TinyStories, the last 200 kept back as a quiz with no Momo in it, and 19,800 others plus 1,000 unique Momo-and-Ella stories that Ella reads three times each, 22,800 stories in training. The tokenizer, the splitter that cuts text into pieces called tokens, is Lesson 01's 4,096-piece one. So are the model code and the training loop. Only two numbers change: the number of blocks stacked (n_layer) and the width of each token's list of numbers (n_embd).
SIZES = { # name -> (n_layer, n_embd); 4 heads, 256-token window, dropout 0 in all
"1m": (2, 128),
"3m": (5, 192),
"6m": (6, 256), # Lesson 01's config
"12m": (6, 384),
}
# per size: seed 42, build GPT, seed 42 again (same batches), 1,500 steps, then measure
for name, (nl, ne) in SIZES.items():
model, cfg, r = train_one(name, nl, ne)
r.update(recite(model)) # 50 prompts x 120 tokens, temperature 0.8, top-k 40, seed 1234
All four models start from the same seed, the number that fixes every random choice so two runs match, and see the same batches in the same order, so size is the only thing that differs. Dropout, which hides some of Momo's thinking at random to discourage copying, stays at zero on purpose.
After training, each model is scored three ways. Quiz loss is how surprised it is by the 200 held-out stories; perplexity is that loss with the logarithm undone, roughly the number of words it is still choosing between. Train loss is the same score on lesson text. And for recitation, the model gets the first 10 tokens of each of 50 training Momo stories, writes 120 more, and I find the longest stretch of that text that appears character for character inside any of the 1,000 training Momo stories. That longest stretch is the run. A 100-character run is about three of the short Momo sentences (they average 32 characters), a quarter of an average 417-character story.
# the longest stretch of a generation that sits inside any training Momo story
sam = SuffixAutomaton("\x00".join(momo)) # \x00 stops a run crossing from one story into the next
run = sam.longest_common(generated_text) # length in characters
The run and its numbers
The "1M" model has 953,856 parameters and the "12M" one 12,318,720. The 6M row lands at 9.0, Lesson 01's figure (its 20-batch estimate; the published weights score 9.09 exactly, as Lesson 04 measures).
| Size | Blocks x width | Parameters | Train loss | Quiz loss | Quiz perplexity | Loss on Momo stories | Train minutes (Mac, this run) |
|---|---|---|---|---|---|---|---|
| 1M | 2 x 128 | 953,856 | 2.836 | 2.773 | 16.0 | 0.554 | 0.55 |
| 3M | 5 x 192 | 3,060,288 | 2.451 | 2.382 | 10.8 | 0.247 | 1.70 |
| 6M | 6 x 256 | 5,853,184 | 2.247 | 2.195 | 9.0 | 0.206 | 3.79 |
| 12M | 6 x 384 | 12,318,720 | 2.024 | 2.008 | 7.4 | 0.185 | 6.33 |
Quiz perplexity falls from 16.0 to 7.4 as the model grows about thirteenfold. Loss on the Momo stories themselves, which every model trained on, drops from 0.554 to 0.185.
| Size | Median longest copied run (characters) | Max run | Generations with a run of 100 or more |
|---|---|---|---|
| 1M | 62 | 126 | 6% (3 of 50) |
| 3M | 211.5 | 361 | 98% (49 of 50) |
| 6M | 265 | 389 | 100% |
| 12M | 269 | 426 | 100% |
The smallest model mostly composes: its median longest copied run is 62 characters and only 3 of 50 generations pass 100. At 3.06 million parameters that flips: 49 of 50 generations contain a run of 100 characters or more, the median is 211.5, and the longest is 361, most of an average story. At 5.85 million and 12.3 million the medians are 265 and 269 and every generation passes 100. The first sample from the two larger models is the same 266-character paragraph, which occurs word for word in three different training stories: a run of stock sentences, exactly the limit the next section describes. The 1M model's samples wander and mix sentences into things like "tell a soft pink banana that".
The check number is 3M. The smallest size whose median verbatim run crosses 100 characters is the 3,060,288-parameter model.
What it means, honestly
Start with what the picture does not show. Train loss (random training batches, which include the Momo stories and so are if anything flattered) sits above quiz loss at every size, the opposite of an overfitting gap; I do not know why the quiz reads easier. Train against quiz alone would never show the larger models recite; the Momo-loss column and the recitation table are where it shows.
Second, this is Momo's friendliest possible case for memorising. The 1,000 Momo stories are generated from small pools of stock sentences and read three times each, so copying is a very cheap way to lower the loss. My run metric cannot tell one story learned by heart from a string of stock sentences shared by many stories, and I did not measure how many generations equal one whole story. Treat "copied run" as a measure of how much of the output is lifted from the training set, not as proof of story-by-story recall. Research such as Quantifying Memorization Across Neural Language Models uses far larger models and messier data; nothing here predicts their thresholds.
Third, the threshold. I trained four sizes, so I can only say that recitation switched on between 0.95M and 3.06M parameters, not where. The 6M and 12M medians differ by 4 characters, which is noise. One seed and 50 prompts means no error bars. Fourth, no quiz story is in training or mentions Momo, though 22 share a stock 20-word TinyStories opening with a training story.
Run it yourself
- Colab. Open Google Colab, upload
lesson_03.ipynb, choose the T4 runtime, run all. It loads Lesson 01's tokenizer from the Hub. - Your own machine. Python 3.10 or newer with torch, transformers, datasets, tokenizers, huggingface_hub and matplotlib. On an Apple Silicon Mac the whole notebook took 10 to 13 minutes.
What you have now: four Momos, 0.95M to 12.3M, and the check number 3M, where median copied runs pass 100 characters. Stage A is done. Next, Stage B: Lesson 04 takes the public Lesson 01 weights, teaches them a new character, and measures what is forgotten.
Check yourself
The number to reproduce: the smallest size whose median longest copied run passes 100 characters. With these seeds it is the 3.06M-parameter model, with a median of 211.5 characters; the 0.95M model's is 62. Medians may shift a few characters between machines; the jump between the first two sizes should stay large.
- Why can the four quiz losses be compared directly?
- Train loss sits above quiz loss at every size. Does that mean the models are not memorising anything? What does the recitation table say instead?
- The Momo stories are built from stock sentences. How would the copied-run numbers change if they were 1,000 genuinely different stories, and which of this post's claims would you stop trusting?
Questions people ask
Why do small language models memorize their training data?
Copying is the cheapest way to lower loss on text seen more than once. The Momo stories are read three times; here recitation switched on between 0.95M and 3.06M parameters.
Is memorisation the same as overfitting?
Not exactly. Overfitting is a train-versus-held-out gap; here it pointed the other way, yet the larger models recite. Memorising a repeated corner of the data need not show in overall curves.
Does this tell me about GPT-scale models?
No. It is four tiny models, a templated corpus and one seed. It shows a method, not where thresholds sit for large models.
References & Citations
- singhpratech/emc2-lesson-03-four-sizes on the Hugging Face Hub (private until reviewed): notebook, RESULTS.md, four sets of weights. One run on 11 October 2026 on an Apple Silicon Mac; the tokenizer is Lesson 01's.
- e=mc² Lesson 01: Tiny Storyteller, whose data pipeline, tokenizer, model and training loop this lesson reuses.
- Eldan, R. and Li, Y. (2023). TinyStories: How Small Can Language Models Be and Still Speak Coherent English? The dataset is roneneldan/TinyStories (CDLA-Sharing-1.0).
- Carlini, N. et al. (2022). Quantifying Memorization Across Neural Language Models.
- All numbers are from that one run; single seed, 50 prompts per size, not timed on a T4.
Subscribe to new posts from theaivibe.org