How Attention Works: Read Tiny Storyteller's Own Attention Maps

Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.
Attention is how a language model lets each word look back at the words before it, and in this lesson you read it off a real model. I open the 5.9-million-parameter Tiny Storyteller from Lesson 01, switch on a flag that records who looked at whom, and look at all 24 attention heads. I plot some, measure all of them over 200 unseen stories, and find the head that puts the most attention on a story's main character's name (22% of its attention, 4.3 times an even spread).
Where you are: Lesson 01 left you a trained Tiny Storyteller on the Hub, 5.9 million dials, quiz perplexity 9, and a 61-word recitation it could not explain. Tonight nothing is trained. I open those same weights and read what the 24 attention heads look at, the first step toward understanding both numbers.
Measured on the 200 held-out quiz stories from Lesson 01, 11 October 2026, on an Apple Silicon Mac. The whole notebook runs in 8 to 20 seconds there.
Momo and Ella, again
Momo the monkey is the model, and his dials are already set from Lesson 01. Ella the elephant is the teacher and the memory: she reads him the story, and she keeps a shelf of everything already said. The story is told in a circle; each word sits in it and is allowed to look at the words that came before.
| Term | Momo's picture | The real thing |
|---|---|---|
| Token | A bite of story | A piece of text with a number on it; 4,096 possible |
| Attention | Looking around the story circle | Each token blends in earlier tokens by relevance |
| Head | One of Momo's four ways of looking in a block | One independent set of query, key and value, 64 numbers wide |
| Query | The question a word asks | A list of 64 numbers made from the word |
| Key | The card a word holds up | A list of 64 numbers made from the word |
| Value | What a word hands over if picked | A list of 64 numbers made from the word |
| Softmax | Turning scores into shares of 100% | Exponentiate the scores, divide by the total |
| Causal mask | No peeking at words not yet said | Scores for later positions set to minus infinity |
| KV cache | Ella's shelf | Stored keys and values of earlier tokens |
| Attention map | The table of who looked at whom | A grid of softmax shares, one row per looking token |
What one head computes
An attention head is one way of looking around the circle. Momo has four ways in each of his six blocks (one attention layer plus its feed-forward layer) (24 in all). A head gives every word three short lists of numbers, 64 numbers each here (in Lesson 01 all three were cards; here I split them so the matching step is visible).
- A query is the question a word asks. When "He" comes up, the question is roughly: who is "he"?
- A key is the card a word holds up. "Momo" holds up a card that says: I am a name.
- A value is what a word hands over if it is picked.
Suppose I am the word "He". The head scores every earlier word by how well its key matches my query (a dot product, multiply and add). A softmax turns the scores into shares that add to 100%: exponentiate each score and divide by the total. Then I take the share-weighted blend of the values. Those shares are the attention probabilities, and a table of them for a whole story is an attention map.
Where the story and the mechanics part ways: nothing is asked in words. A query is 64 numbers produced by a matrix multiplication. The label "looking for a name" is something I read into a head afterwards; Momo was only ever trained to guess the next word.
The triangle, and Ella's shelf
Momo must never peek at words that come after the one he is guessing. The causal mask enforces it: before the softmax, every score for a later word is set to minus infinity, so its share is exactly zero. The map of a whole story is therefore a lower triangle. The dark upper-right corner of every heatmap is the mask, not a choice the head made.
The mask also explains Ella's shelf, the KV cache (key and value cache). Earlier words cannot see later ones, so an earlier word's key and value never change. When Momo writes a story one token at a time, Ella files each token's key and value on the shelf; the next token computes only its own query, key and value and reads the shelf. In the notebook, earlier keys and values computed from the shorter prefix alone equal the full run's exactly, and the shelf's answer matches recomputing everything within 0.0000003. Writing 200 tokens costs 200 key and value computations instead of 20,100. It is a speed trick, not a memory of the story: what Momo knows lives in the dials, not on the shelf.
What the notebook does
It copies the GPT class from Lesson 01 and adds one flag. With the flag on, each block computes attention the long way and keeps the probabilities; off, it uses PyTorch's built-in fast attention routine (the kernel) as before. The two agree within 0.00002 in the raw output scores (logits), and every attention row sums to 1.
Then it plots the causal-mask triangle, all 24 heads on one story, and a clickable toy widget (not Momo's weights). The measurement uses the last 200 of Lesson 01's 20,000 TinyStories, the quiz Momo never trained on. Leakage check: none of the 200 is identical to a training story or mentions Momo, and a quiz story's 8-word runs found in the 19,800 training stories average 8.7% (maximum 30%), reading the split as Lesson 01 defines it; I did not retrain.
Finding the name-tracking head
The definition is crude, and five of the 200 stories were cut to the 256-token window. A name is a capitalised word whose lowercase form appears nowhere in the 200 quiz stories, which discards "She" and "The" and keeps "Lily" and "Tim". The main character is the most frequent such name in the story, with at least two mentions. 159 of the 200 stories have one; the other 41 are skipped. For every (block, head) I take every non-name position after the first mention and add up the attention it puts on the name's tokens, average within a story, then across the 159.
Name mass is the share of a head's attention that lands on the name's tokens. Against a baseline of what an even spread would give, 0.052, here are the top six of 24.
| Rank | Head (block, head) | Mean attention on the name | Against even spread (0.052) |
|---|---|---|---|
| 1 | L5H3 | 0.221 | 4.3x |
| 2 | L4H4 | 0.199 | 3.9x |
| 3 | L6H2 | 0.178 | 3.4x |
| 4 | L4H3 | 0.122 | 2.4x |
| 5 | L4H2 | 0.079 | 1.5x |
| 6 | L1H2 | 0.060 | 1.2x |
The winner is block 5, head 3 (L5H3), with 0.221 of its attention on the name, 4.3 times an even spread. L4H4 and L6H2 are close behind. At the other end, 17 of the 24 heads put less on the name than an even spread would. In the story I plotted, L5H3 also lights up on "kid" and "He", which refer to the same character; my measure counts only name tokens.
What it means, honestly
Three honest limits. First, L5H3 puts about 22% of its attention on the name, not all of it; this is a tendency, not a switch. Second, attention is where a head looks, not proof of what the model relies on. I did not switch any head off. Third, name position is not controlled: a name early in a story may draw attention partly for being early.
The disappointing part: previous-token heads, with more than half their attention on the word just before. The count is 0 of 24. The highest is L1H1 at 0.146; counting how often the previous word is the single most-attended one, the best head reaches 0.369 of queries. L1H1 puts 0.140 on the word just before and 0.389 on the last four together (first 100 quiz stories, as is the 0.369 figure): smeared over the last few words, not fixed on one. No head leans mostly on the end-of-story marker (the token Lesson 01's tokenizer puts after each story; it also opens the next) either; the highest is 0.065.
Run it yourself
- Colab. Open Google Colab, upload lesson_02.ipynb from the lesson repo, set the runtime to T4, run all.
- Your machine. Python 3.10+ with torch, transformers, datasets, tokenizers and matplotlib. It loads the public Lesson 01 weights.
What you have now: an attention-map notebook over the Lesson 01 weights and one head, L5H3, that puts 22% of its attention on the main character's name. Next, Lesson 03 goes back to the recitation: the same recipe at four sizes, to find where copying begins.
Check yourself
The number to reproduce: the most name-attending head is block 5, head 3, with mean name mass about 0.221 over 159 of the 200 stories, against an even-spread baseline of 0.052.
- Why is the upper-right of every attention map exactly zero?
- Why can Ella's shelf store keys and values but never need to update them?
- L5H3 sends 22% of its attention to the name. Does that prove the model uses the name to write the next word? What experiment would tell you?
Part of e=mc²; attention comes from Attention Is All You Need.
Questions people ask
What is the KV cache and why do language models use it?
Each new token needs the keys and values of every earlier token, which never change once computed. The KV cache stores them so each step computes only the newest token's.
What is the causal mask in a transformer?
A rule that stops a token looking at later tokens: their scores are set to minus infinity before the softmax, so their attention is exactly zero.
Does a head that attends to a name mean the model understands the character?
No. It means that head sends about 22% of its attention to the name tokens. I did not switch heads off, so this describes where attention goes, not what the model relies on.
References & Citations
- singhpratech/emc2-lesson-02-attention on the Hugging Face Hub (private until reviewed): lesson_02.ipynb, RESULTS.md, results.json, figures. Run 11 October 2026 on an Apple Silicon Mac.
- singhpratech/tiny-storyteller: the Lesson 01 weights, config and tokenizer used here.
- Eldan, R. and Li, Y. (2023). TinyStories; dataset roneneldan/TinyStories (CDLA-Sharing-1.0).
- Vaswani, A. et al. (2017). Attention Is All You Need.
- All numbers are from RESULTS.md in the lesson repo: 200 held-out quiz stories, 159 with a main character, name mass measured on non-name positions after the first mention, even-spread baseline 0.0516.
Subscribe to new posts from theaivibe.org