
e=mc²9 min read
How Attention Works: Read Tiny Storyteller's Own Attention Maps
Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.
2 views
Read