Skip to content

How Attention Works: Read Tiny Storyteller's Own Attention Maps

Prateek SinghOctober 11, 20269 min read2 views
How Attention Works: Read Tiny Storyteller's Own Attention Maps

Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.

Attention is how a language model lets each word look back at the words before it, and in this lesson you read it off a real model. I open the 5.9-million-parameter Tiny Storyteller from Lesson 01, switch on a flag that records who looked at whom, and look at all 24 attention heads. I plot some, measure all of them over 200 unseen stories, and find the head that puts the most attention on a story's main character's name (22% of its attention, 4.3 times an even spread).

Where you are: Lesson 01 left you a trained Tiny Storyteller on the Hub, 5.9 million dials, quiz perplexity 9, and a 61-word recitation it could not explain. Tonight nothing is trained. I open those same weights and read what the 24 attention heads look at, the first step toward understanding both numbers.

attention heads read
24
6 blocks of 4 heads, all from the Lesson 01 weights; no training
most name-attending head
L5H3
block 5, head 3: 22% of its attention lands on the main character's name
times the even-spread baseline
4.3x
0.221 against 0.052, what spreading attention evenly would give
previous-token heads
0 of 24
none put more than half their attention on the word just before; the highest is 0.146

Measured on the 200 held-out quiz stories from Lesson 01, 11 October 2026, on an Apple Silicon Mac. The whole notebook runs in 8 to 20 seconds there.

Momo and Ella, again

Momo the monkey is the model, and his dials are already set from Lesson 01. Ella the elephant is the teacher and the memory: she reads him the story, and she keeps a shelf of everything already said. The story is told in a circle; each word sits in it and is allowed to look at the words that came before.

TermMomo's pictureThe real thing
TokenA bite of storyA piece of text with a number on it; 4,096 possible
AttentionLooking around the story circleEach token blends in earlier tokens by relevance
HeadOne of Momo's four ways of looking in a blockOne independent set of query, key and value, 64 numbers wide
QueryThe question a word asksA list of 64 numbers made from the word
KeyThe card a word holds upA list of 64 numbers made from the word
ValueWhat a word hands over if pickedA list of 64 numbers made from the word
SoftmaxTurning scores into shares of 100%Exponentiate the scores, divide by the total
Causal maskNo peeking at words not yet saidScores for later positions set to minus infinity
KV cacheElla's shelfStored keys and values of earlier tokens
Attention mapThe table of who looked at whomA grid of softmax shares, one row per looking token
The ten words this lesson uses; the notebook carries the same table.

What one head computes

An attention head is one way of looking around the circle. Momo has four ways in each of his six blocks (one attention layer plus its feed-forward layer) (24 in all). A head gives every word three short lists of numbers, 64 numbers each here (in Lesson 01 all three were cards; here I split them so the matching step is visible).

  • A query is the question a word asks. When "He" comes up, the question is roughly: who is "he"?
  • A key is the card a word holds up. "Momo" holds up a card that says: I am a name.
  • A value is what a word hands over if it is picked.

Suppose I am the word "He". The head scores every earlier word by how well its key matches my query (a dot product, multiply and add). A softmax turns the scores into shares that add to 100%: exponentiate each score and divide by the total. Then I take the share-weighted blend of the values. Those shares are the attention probabilities, and a table of them for a whole story is an attention map.

Where the story and the mechanics part ways: nothing is asked in words. A query is 64 numbers produced by a matrix multiplication. The label "looking for a name" is something I read into a head afterwards; Momo was only ever trained to guess the next word.

The triangle, and Ella's shelf

Momo must never peek at words that come after the one he is guessing. The causal mask enforces it: before the softmax, every score for a later word is set to minus infinity, so its share is exactly zero. The map of a whole story is therefore a lower triangle. The dark upper-right corner of every heatmap is the mask, not a choice the head made.

The mask also explains Ella's shelf, the KV cache (key and value cache). Earlier words cannot see later ones, so an earlier word's key and value never change. When Momo writes a story one token at a time, Ella files each token's key and value on the shelf; the next token computes only its own query, key and value and reads the shelf. In the notebook, earlier keys and values computed from the shorter prefix alone equal the full run's exactly, and the shelf's answer matches recomputing everything within 0.0000003. Writing 200 tokens costs 200 key and value computations instead of 20,100. It is a speed trick, not a memory of the story: what Momo knows lives in the dials, not on the shelf.

What the notebook does

It copies the GPT class from Lesson 01 and adds one flag. With the flag on, each block computes attention the long way and keeps the probabilities; off, it uses PyTorch's built-in fast attention routine (the kernel) as before. The two agree within 0.00002 in the raw output scores (logits), and every attention row sums to 1.

Then it plots the causal-mask triangle, all 24 heads on one story, and a clickable toy widget (not Momo's weights). The measurement uses the last 200 of Lesson 01's 20,000 TinyStories, the quiz Momo never trained on. Leakage check: none of the 200 is identical to a training story or mentions Momo, and a quiz story's 8-word runs found in the 19,800 training stories average 8.7% (maximum 30%), reading the split as Lesson 01 defines it; I did not retrain.

Finding the name-tracking head

The definition is crude, and five of the 200 stories were cut to the 256-token window. A name is a capitalised word whose lowercase form appears nowhere in the 200 quiz stories, which discards "She" and "The" and keeps "Lily" and "Tim". The main character is the most frequent such name in the story, with at least two mentions. 159 of the 200 stories have one; the other 41 are skipped. For every (block, head) I take every non-name position after the first mention and add up the attention it puts on the name's tokens, average within a story, then across the 159.

Name mass is the share of a head's attention that lands on the name's tokens. Against a baseline of what an even spread would give, 0.052, here are the top six of 24.

RankHead (block, head)Mean attention on the nameAgainst even spread (0.052)
1L5H30.2214.3x
2L4H40.1993.9x
3L6H20.1783.4x
4L4H30.1222.4x
5L4H20.0791.5x
6L1H20.0601.2x
The six heads with the most attention on the main character's name, out of 24. Mean over 159 stories.

The winner is block 5, head 3 (L5H3), with 0.221 of its attention on the name, 4.3 times an even spread. L4H4 and L6H2 are close behind. At the other end, 17 of the 24 heads put less on the name than an even spread would. In the story I plotted, L5H3 also lights up on "kid" and "He", which refer to the same character; my measure counts only name tokens.

What it means, honestly

Three honest limits. First, L5H3 puts about 22% of its attention on the name, not all of it; this is a tendency, not a switch. Second, attention is where a head looks, not proof of what the model relies on. I did not switch any head off. Third, name position is not controlled: a name early in a story may draw attention partly for being early.

The disappointing part: previous-token heads, with more than half their attention on the word just before. The count is 0 of 24. The highest is L1H1 at 0.146; counting how often the previous word is the single most-attended one, the best head reaches 0.369 of queries. L1H1 puts 0.140 on the word just before and 0.389 on the last four together (first 100 quiz stories, as is the 0.369 figure): smeared over the last few words, not fixed on one. No head leans mostly on the end-of-story marker (the token Lesson 01's tokenizer puts after each story; it also opens the next) either; the highest is 0.065.

Run it yourself

What you have now: an attention-map notebook over the Lesson 01 weights and one head, L5H3, that puts 22% of its attention on the main character's name. Next, Lesson 03 goes back to the recitation: the same recipe at four sizes, to find where copying begins.

Check yourself

The number to reproduce: the most name-attending head is block 5, head 3, with mean name mass about 0.221 over 159 of the 200 stories, against an even-spread baseline of 0.052.

  1. Why is the upper-right of every attention map exactly zero?
  2. Why can Ella's shelf store keys and values but never need to update them?
  3. L5H3 sends 22% of its attention to the name. Does that prove the model uses the name to write the next word? What experiment would tell you?

Part of e=mc²; attention comes from Attention Is All You Need.

Questions people ask

What is the KV cache and why do language models use it?

Each new token needs the keys and values of every earlier token, which never change once computed. The KV cache stores them so each step computes only the newest token's.

What is the causal mask in a transformer?

A rule that stops a token looking at later tokens: their scores are set to minus infinity before the softmax, so their attention is exactly zero.

Does a head that attends to a name mean the model understands the character?

No. It means that head sends about 22% of its attention to the name tokens. I did not switch heads off, so this describes where attention goes, not what the model relies on.

References & Citations

  • singhpratech/emc2-lesson-02-attention on the Hugging Face Hub (private until reviewed): lesson_02.ipynb, RESULTS.md, results.json, figures. Run 11 October 2026 on an Apple Silicon Mac.
  • singhpratech/tiny-storyteller: the Lesson 01 weights, config and tokenizer used here.
  • Eldan, R. and Li, Y. (2023). TinyStories; dataset roneneldan/TinyStories (CDLA-Sharing-1.0).
  • Vaswani, A. et al. (2017). Attention Is All You Need.
  • All numbers are from RESULTS.md in the lesson repo: 200 held-out quiz stories, 159 with a main character, name mass measured on non-name positions after the first mention, even-spread baseline 0.0516.

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article