Skip to content

Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Prateek SinghOctober 11, 20269 min read2 views
Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Lesson 01: a 5.9-million-parameter GPT trained from random weights on children's stories: 1,500 steps in 1.2 minutes of training on a free Colab GPU (the whole notebook is about 5 minutes). Every piece explained through Momo the monkey and Ella the elephant, with the real numbers and an honest look at what it wrote.

This is the first lesson of e=mc², and it builds a language model from nothing. No pretrained weights and nothing hidden: a tokenizer learned from the data, a GPT-style transformer in plain PyTorch, and a training loop you can watch. It is called Tiny Storyteller because its one job is to continue a bedtime story for a young child, which it learns in 1,500 steps in 1.2 minutes of training on a free Google Colab GPU (the whole Colab notebook is about 5 minutes).

This is where e=mc² starts, so there is nothing to build on yet: no model, no tokenizer, no number. By the end you will have a 5.9-million-parameter storyteller on the Hugging Face Hub and one number to remember, its quiz perplexity of 9. Every later lesson opens that model or copies its recipe.

parameters, all starting random
5.9M
6 blocks, 4 heads, 256-wide; vocabulary of 4,096 pieces
training steps
1,500
batch of 32 windows, 256 tokens each
training time on a free Colab T4
1.2 min
about 5 minutes end to end; about 2 minutes to train on an Apple Silicon Mac
perplexity on 200 unseen stories
9
runs land around 9 to 10; random guessing would be 4,096

Final run, 10 October 2026, identical on Colab's T4 and on an Apple Silicon Mac. About 23,000 training stories (20,000 TinyStories, less 200 for the quiz, plus 1,000 Momo stories read three times); the quiz stories were never trained on.

The model, tokenizer, model card and the notebook that builds it live on the Hugging Face Hub. The notebook is the lesson proper: 66 cells and ten widgets.

Two characters, so nothing has to be re-learned

The whole build is told through two characters. Momo is a young monkey learning to tell stories. Inside his head are millions of tiny dials, all pointing somewhere random, so his first stories are nonsense. Ella is an elephant, and elephants never forget. Every night she reads Momo a story one word at a time, stops before each word, lets him guess, and tells him how wrong he was. She also keeps a shelf of everything already said so nobody works it out twice.

In grown-up words: Momo is the model, his dials are the weights, Ella's stories are the training data, her "how wrong were you" is the loss, and her shelf is the KV cache. The ten words below are the ones this post needs.

WordIn Momo's worldGrown-up meaning
ModelMomoA pile of dials plus a recipe for using them
Weights, parametersMomo's dialsThe numbers training adjusts; 5.9 million here
Training dataElla's bedtime storiesThe text the model learns from
Token, tokenizerA bite, the biterA piece of text, and the splitter that makes the pieces
EmbeddingA spot on Momo's meaning mapA list of 256 numbers standing for one token
AttentionLooking around the story circleEach token blends in earlier tokens by relevance
LossHow wrong was that guessThe score training pushes down
GradientThe slope under Momo's feetWhich way to turn each dial, and how far
KV cacheElla's shelfStored keys and values of past tokens; nothing is redone
PerplexityBetween how many words is Momo still choosingexp(loss); 4,096 is random, 9 to 10 is real sentences
The ten words this post uses; the notebook decodes about forty.

What a language model does, in one sentence

Given the words so far, guess the next one. A story is that guess repeated; a chat answer is the same, with your question as the words so far.

When Momo continues "Once upon a time, a little", four things happen. The sentence is bitten into tokens, pieces of text with a number on each; the vocabulary has 4,096. Each token becomes a spot on a meaning map, a list of 256 numbers, where similar tokens sit close together. Then every token looks back at the tokens before it and decides how much to listen to each: "capital" finds "France" and pulls some of its meaning across. That is attention, from the 2017 paper Attention Is All You Need, and it runs six times in a row, four heads each. Finally the last token's sharpened meaning is scored against every token in the vocabulary and one is picked.

What the notebook builds, part by part

The lesson adds one idea at a time; the KV cache, Ella's shelf, is checked against the from-scratch answer.

PartWhat you buildWhat you learn
0Momo, Ella, a word listEvery idea before any code
1Tensors and gradientsBaskets of numbers; which way to turn a dial
2Data: TinyStories, plus 1,000 Momo stories200 stories hidden as the quiz
3A bigram modelLetter soup, but the loss falls
4A tokenizer (byte-pair encoding)4,096 pieces learned from the stories
5Attention, by handQuery, key, value; the mask; the KV cache
6A tiny GPTEmbeddings, 6 blocks, a scoring head
7TrainingCross-entropy, AdamW, warm-up and cosine decay; 1,500 steps
8Telling a storyTemperature and top-k; what it wrote, honestly
9Publish to the Hugging Face HubWeights, config, tokenizer, model card
10Load it back laterRebuild the model on any machine
The notebook's road map; most parts have a widget that shows the idea first.

The model is a standard small GPT, written out line by line in Part 6: each token gets its spot on the meaning map plus a second spot for its seat number, climbs six blocks of attention and a small think-alone network, and ends at a scoring tray that reuses the meaning map (weight tying) so fewer dials are needed. You see the three attention cards and the triangle mask that stops a token peeking ahead.

Training, and what the numbers mean

Every training step is the same five moves. Ella picks a batch, 32 windows of 256 tokens, each shifted by one token as the right answers. Momo guesses the next token at every position. Ella scores the guesses with cross-entropy: how surprised was he by the real next word? The blame runs backwards through the tape PyTorch recorded, and every dial gets its gradient. AdamW, the optimizer, turns the dials. The size of each turn, the learning rate, starts tiny, reaches full stride, then shrinks along a cosine curve. That is 1,500 steps in 1.2 minutes of training on the T4 (the whole Colab notebook is about 5 minutes), and about 2 minutes on the Mac.

The number to watch is perplexity, the loss with the logarithm undone: roughly, between how many tokens is the model still choosing. Random guessing is 4,096. This run finished at 9 on the quiz; the runs before it landed around 9 to 10. The quiz is the part I care most about: the last 200 TinyStories are held out before training, with no Momo stories in them, so a story the model saw three times can never flatter the score. An earlier draft got that wrong, and fixing it is why the number is honest now.

What it wrote, honestly

Given "Lily had a red balloon.", one run produced this:

She ran to her mom and gave her a big bite. They both ate the balloon and had the balloon. It was shiny and yellow. They ran to the park to play with.

Grammatical, kid-flavoured, slightly unhinged: what 5.9 million parameters and two minutes buy. The stories without Momo wander, forget who is speaking and sometimes rename a character halfway. The Momo stories are a different failure. Given "Once upon a time, a little monkey", one run opened with a bad dream and then wrote this:

Momo was sad because he could not swim like Bo the bear cub. He sat in the tall grass with his feet in the water. Ella said, "Try again, little one. Every time you try, you learn a little more." She held him up with her trunk. Momo kicked and kicked. Soon he could swim all by himself. The end.

Those 61 words are, character for character, the second half of one of the thousand training stories, which the model read three times. That is memorising, not inventing; a few thousand repeats are easy for 5.9 million dials to store. The fix is more unique stories and fewer repeats.

Run it yourself

The playbook has the long version.

  • Colab, no install. Upload the notebook to Google Colab, set the runtime to T4 GPU, run all: about five minutes. Part 8 is where you type your own prompts.
  • Your own machine. Python 3.10+ with torch, transformers, datasets and tokenizers; about four minutes on an Apple Silicon Mac, slower on a CPU.
  • Skip training. Load the published weights from the Hub:
from huggingface_hub import hf_hub_download
from transformers import PreTrainedTokenizerFast
import json, torch

# Config and GPT are from Part 6 of the notebook, tell_story() from Part 8
cfg   = json.load(open(hf_hub_download("singhpratech/tiny-storyteller", "config.json")))
state = torch.load(hf_hub_download("singhpratech/tiny-storyteller", "model.pt"), map_location="cpu")
tok   = PreTrainedTokenizerFast.from_pretrained("singhpratech/tiny-storyteller")

model = GPT(Config(**cfg)); model.load_state_dict(state); model.eval()
print(tell_story("Once upon a time, a little monkey", temperature=0.8))

Start prompts the way the stories start ("Once upon a time, a little monkey"); temperature 0.8 is the default. Publishing your own copy needs a Hugging Face write token as a Colab secret.

Why this is lesson one

This vertical assumes nothing. The recipe is the one the large models use, with more stories, more blocks and more GPUs; once you have turned 5.9 million dials yourself and watched perplexity fall from 4,096 to 9, the big numbers stop being magic and start being engineering.

What you have now: singhpratech/tiny-storyteller on the Hub, 5.9M dials, quiz perplexity 9, and one unexplained 61-word recitation. Next, Lesson 02 opens these exact weights, no training, and reads what the 24 attention heads look at.

Questions people ask

Is a 5.9-million-parameter model really an LLM?

It is a language model with the same recipe as the large ones: byte-pair tokenizer, embeddings, stacked attention blocks, a next-token head, cross-entropy and AdamW. It is not large; the recipe does not change as it scales.

Why children's stories?

TinyStories is a public dataset of short stories in the vocabulary a young child knows, built to show that very small models can write coherent English if the language is simple. This size trained on Wikipedia would write nothing readable.

Can I train it on my own text?

Yes. The notebook reads a list of strings; replace the TinyStories slice with your own text and the same code learns that style. Expect it to recite passages it has seen many times, as it does with the Momo stories.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article