Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Lesson 01: a 5.9-million-parameter GPT trained from random weights on children's stories: 1,500 steps in 1.2 minutes of training on a free Colab GPU (the whole notebook is about 5 minutes). Every piece explained through Momo the monkey and Ella the elephant, with the real numbers and an honest look at what it wrote.
This is the first lesson of e=mc², and it builds a language model from nothing. No pretrained weights and nothing hidden: a tokenizer learned from the data, a GPT-style transformer in plain PyTorch, and a training loop you can watch. It is called Tiny Storyteller because its one job is to continue a bedtime story for a young child, which it learns in 1,500 steps in 1.2 minutes of training on a free Google Colab GPU (the whole Colab notebook is about 5 minutes).
This is where e=mc² starts, so there is nothing to build on yet: no model, no tokenizer, no number. By the end you will have a 5.9-million-parameter storyteller on the Hugging Face Hub and one number to remember, its quiz perplexity of 9. Every later lesson opens that model or copies its recipe.
Final run, 10 October 2026, identical on Colab's T4 and on an Apple Silicon Mac. About 23,000 training stories (20,000 TinyStories, less 200 for the quiz, plus 1,000 Momo stories read three times); the quiz stories were never trained on.
The model, tokenizer, model card and the notebook that builds it live on the Hugging Face Hub. The notebook is the lesson proper: 66 cells and ten widgets.
Two characters, so nothing has to be re-learned
The whole build is told through two characters. Momo is a young monkey learning to tell stories. Inside his head are millions of tiny dials, all pointing somewhere random, so his first stories are nonsense. Ella is an elephant, and elephants never forget. Every night she reads Momo a story one word at a time, stops before each word, lets him guess, and tells him how wrong he was. She also keeps a shelf of everything already said so nobody works it out twice.
In grown-up words: Momo is the model, his dials are the weights, Ella's stories are the training data, her "how wrong were you" is the loss, and her shelf is the KV cache. The ten words below are the ones this post needs.
| Word | In Momo's world | Grown-up meaning |
|---|---|---|
| Model | Momo | A pile of dials plus a recipe for using them |
| Weights, parameters | Momo's dials | The numbers training adjusts; 5.9 million here |
| Training data | Ella's bedtime stories | The text the model learns from |
| Token, tokenizer | A bite, the biter | A piece of text, and the splitter that makes the pieces |
| Embedding | A spot on Momo's meaning map | A list of 256 numbers standing for one token |
| Attention | Looking around the story circle | Each token blends in earlier tokens by relevance |
| Loss | How wrong was that guess | The score training pushes down |
| Gradient | The slope under Momo's feet | Which way to turn each dial, and how far |
| KV cache | Ella's shelf | Stored keys and values of past tokens; nothing is redone |
| Perplexity | Between how many words is Momo still choosing | exp(loss); 4,096 is random, 9 to 10 is real sentences |
What a language model does, in one sentence
Given the words so far, guess the next one. A story is that guess repeated; a chat answer is the same, with your question as the words so far.
When Momo continues "Once upon a time, a little", four things happen. The sentence is bitten into tokens, pieces of text with a number on each; the vocabulary has 4,096. Each token becomes a spot on a meaning map, a list of 256 numbers, where similar tokens sit close together. Then every token looks back at the tokens before it and decides how much to listen to each: "capital" finds "France" and pulls some of its meaning across. That is attention, from the 2017 paper Attention Is All You Need, and it runs six times in a row, four heads each. Finally the last token's sharpened meaning is scored against every token in the vocabulary and one is picked.
What the notebook builds, part by part
The lesson adds one idea at a time; the KV cache, Ella's shelf, is checked against the from-scratch answer.
| Part | What you build | What you learn |
|---|---|---|
| 0 | Momo, Ella, a word list | Every idea before any code |
| 1 | Tensors and gradients | Baskets of numbers; which way to turn a dial |
| 2 | Data: TinyStories, plus 1,000 Momo stories | 200 stories hidden as the quiz |
| 3 | A bigram model | Letter soup, but the loss falls |
| 4 | A tokenizer (byte-pair encoding) | 4,096 pieces learned from the stories |
| 5 | Attention, by hand | Query, key, value; the mask; the KV cache |
| 6 | A tiny GPT | Embeddings, 6 blocks, a scoring head |
| 7 | Training | Cross-entropy, AdamW, warm-up and cosine decay; 1,500 steps |
| 8 | Telling a story | Temperature and top-k; what it wrote, honestly |
| 9 | Publish to the Hugging Face Hub | Weights, config, tokenizer, model card |
| 10 | Load it back later | Rebuild the model on any machine |
The model is a standard small GPT, written out line by line in Part 6: each token gets its spot on the meaning map plus a second spot for its seat number, climbs six blocks of attention and a small think-alone network, and ends at a scoring tray that reuses the meaning map (weight tying) so fewer dials are needed. You see the three attention cards and the triangle mask that stops a token peeking ahead.
Training, and what the numbers mean
Every training step is the same five moves. Ella picks a batch, 32 windows of 256 tokens, each shifted by one token as the right answers. Momo guesses the next token at every position. Ella scores the guesses with cross-entropy: how surprised was he by the real next word? The blame runs backwards through the tape PyTorch recorded, and every dial gets its gradient. AdamW, the optimizer, turns the dials. The size of each turn, the learning rate, starts tiny, reaches full stride, then shrinks along a cosine curve. That is 1,500 steps in 1.2 minutes of training on the T4 (the whole Colab notebook is about 5 minutes), and about 2 minutes on the Mac.
The number to watch is perplexity, the loss with the logarithm undone: roughly, between how many tokens is the model still choosing. Random guessing is 4,096. This run finished at 9 on the quiz; the runs before it landed around 9 to 10. The quiz is the part I care most about: the last 200 TinyStories are held out before training, with no Momo stories in them, so a story the model saw three times can never flatter the score. An earlier draft got that wrong, and fixing it is why the number is honest now.
What it wrote, honestly
Given "Lily had a red balloon.", one run produced this:
She ran to her mom and gave her a big bite. They both ate the balloon and had the balloon. It was shiny and yellow. They ran to the park to play with.
Grammatical, kid-flavoured, slightly unhinged: what 5.9 million parameters and two minutes buy. The stories without Momo wander, forget who is speaking and sometimes rename a character halfway. The Momo stories are a different failure. Given "Once upon a time, a little monkey", one run opened with a bad dream and then wrote this:
Momo was sad because he could not swim like Bo the bear cub. He sat in the tall grass with his feet in the water. Ella said, "Try again, little one. Every time you try, you learn a little more." She held him up with her trunk. Momo kicked and kicked. Soon he could swim all by himself. The end.
Those 61 words are, character for character, the second half of one of the thousand training stories, which the model read three times. That is memorising, not inventing; a few thousand repeats are easy for 5.9 million dials to store. The fix is more unique stories and fewer repeats.
Run it yourself
The playbook has the long version.
- Colab, no install. Upload the notebook to Google Colab, set the runtime to T4 GPU, run all: about five minutes. Part 8 is where you type your own prompts.
- Your own machine. Python 3.10+ with torch, transformers, datasets and tokenizers; about four minutes on an Apple Silicon Mac, slower on a CPU.
- Skip training. Load the published weights from the Hub:
from huggingface_hub import hf_hub_download
from transformers import PreTrainedTokenizerFast
import json, torch
# Config and GPT are from Part 6 of the notebook, tell_story() from Part 8
cfg = json.load(open(hf_hub_download("singhpratech/tiny-storyteller", "config.json")))
state = torch.load(hf_hub_download("singhpratech/tiny-storyteller", "model.pt"), map_location="cpu")
tok = PreTrainedTokenizerFast.from_pretrained("singhpratech/tiny-storyteller")
model = GPT(Config(**cfg)); model.load_state_dict(state); model.eval()
print(tell_story("Once upon a time, a little monkey", temperature=0.8))
Start prompts the way the stories start ("Once upon a time, a little monkey"); temperature 0.8 is the default. Publishing your own copy needs a Hugging Face write token as a Colab secret.
Why this is lesson one
This vertical assumes nothing. The recipe is the one the large models use, with more stories, more blocks and more GPUs; once you have turned 5.9 million dials yourself and watched perplexity fall from 4,096 to 9, the big numbers stop being magic and start being engineering.
What you have now: singhpratech/tiny-storyteller on the Hub, 5.9M dials, quiz perplexity 9, and one unexplained 61-word recitation. Next, Lesson 02 opens these exact weights, no training, and reads what the 24 attention heads look at.
Questions people ask
Is a 5.9-million-parameter model really an LLM?
It is a language model with the same recipe as the large ones: byte-pair tokenizer, embeddings, stacked attention blocks, a next-token head, cross-entropy and AdamW. It is not large; the recipe does not change as it scales.
Why children's stories?
TinyStories is a public dataset of short stories in the vocabulary a young child knows, built to show that very small models can write coherent English if the language is simple. This size trained on Wikipedia would write nothing readable.
Can I train it on my own text?
Yes. The notebook reads a list of strings; replace the TinyStories slice with your own text and the same code learns that style. Expect it to recite passages it has seen many times, as it does with the Momo stories.
References & Citations
- singhpratech/tiny-storyteller on the Hugging Face Hub: weights (
model.pt),config.json, tokenizer, model card, tiny_llm_from_scratch.ipynb and PLAYBOOK.md. Final run 10 October 2026. - Eldan, R. and Li, Y. (2023). TinyStories: How Small Can Language Models Be and Still Speak Coherent English? The dataset is roneneldan/TinyStories (CDLA-Sharing-1.0).
- Vaswani, A. et al. (2017). Attention Is All You Need.
- All numbers are from the notebook's final run on Google Colab (T4) and an Apple Silicon Mac, both 10 October 2026; the quiz perplexity is on 200 held-out TinyStories with no generated stories in them. The Lily sample is from one generation at temperature 0.8; the Momo sample was generated from the published weights on 10 October 2026 (seed 1, temperature 0.8, top-k 40) and matched against the generated training stories with a longest-common-run check.
Subscribe to new posts from theaivibe.org