Skip to content

Tag

Hugging Face

6 posts

LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights
e=mc²9 min read

LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights

Lesson 06: LoRA written by hand in a few dozen lines of PyTorch, training 2.11% of a 5.9M-parameter model on the same task as a full fine-tune. The adapter is 0.51 MB instead of 23.44 MB and merges back to within rounding. It was not faster per step, and it followed fewer requests. The real numbers.

2 views
Read
Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower
e=mc²9 min read

Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower

Lesson 05: teach a 5.9M-parameter story model to follow a request, with two special tokens, a chat template and prompt masking. On 50 held-out requests, the share of stories with the right character and place went from 12% to 44%. Here is what it learned and what it did not.

2 views
Read
Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting
e=mc²9 min read

Fine-Tune a Small LLM: Add a New Character and Measure the Forgetting

Lesson 04: fine-tune the 5.9M-parameter Tiny Storyteller on 300 stories about a new character. Its old quiz perplexity jumps from 9.09 to 156. Mixing two old windows into every batch of twenty holds it at 11.75. Real numbers, and what the model got wrong.

2 views
Read
Why Momo Recites: Train Four Sizes and Watch Memorisation Begin
e=mc²9 min read

Why Momo Recites: Train Four Sizes and Watch Memorisation Begin

Lesson 03: the same small language model trained at 0.95M, 3.06M, 5.85M and 12.3M parameters on the same stories. Quiz perplexity falls from 16.0 to 7.4, and verbatim recitation of training stories switches on somewhere between the first and second size. Real numbers, and what they do not prove.

2 views
Read
How Attention Works: Read Tiny Storyteller's Own Attention Maps
e=mc²9 min read

How Attention Works: Read Tiny Storyteller's Own Attention Maps

Lesson 02: open up the 5.9-million-parameter model from Lesson 01 and read what all 24 attention heads look at. Query, key and value in plain words, the causal mask, why the KV cache exists, and the head that puts the most attention on a story's main character's name.

2 views
Read
Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller
e=mc²9 min read

Build an LLM From Scratch on a Free GPU in 5 Minutes: Tiny Storyteller

Lesson 01: a 5.9-million-parameter GPT trained from random weights on children's stories: 1,500 steps in 1.2 minutes of training on a free Colab GPU (the whole notebook is about 5 minutes). Every piece explained through Momo the monkey and Ella the elephant, with the real numbers and an honest look at what it wrote.

2 views
Read