
Why Momo Recites: Train Four Sizes and Watch Memorisation Begin
Lesson 03: the same small language model trained at 0.95M, 3.06M, 5.85M and 12.3M parameters on the same stories. Quiz perplexity falls from 16.0 to 7.4, and verbatim recitation of training stories switches on somewhere between the first and second size. Real numbers, and what they do not prove.

