Instruction Tuning From Scratch: Turn a Story Generator Into a Prompt Follower

Lesson 05: teach a 5.9M-parameter story model to follow a request, with two special tokens, a chat template and prompt masking. On 50 held-out requests, the share of stories with the right character and place went from 12% to 44%. Here is what it learned and what it did not.
Lesson 05 of e=mc² turns a model that continues stories into a model that takes requests. That step, instruction tuning (fine-tuning on request-and-answer pairs), is what separates a base model from an assistant, and it is small enough to do by hand: the pairs, a chat template (the fixed order requests and answers are written in) with two new tokens, and one line that stops the model being graded on the request. Done on the 5.9-million-parameter storyteller from Lesson 04, it raised the share of held-out requests whose story had the right character and place from 12% to 44%.
Where you are: Lesson 04 gave you tiny-storyteller-v2, Momo who knows Zuri and still scores 11.75 on the old quiz because two old windows rode in every batch. He can only continue a story. Tonight he learns to take a request, the step that turns a base model into an assistant, with two new tokens and one masked loss.
Run of 11 October 2026 on an Apple Silicon Mac. Starting point: tiny-storyteller-v2 from Lesson 04, 5.9M parameters.
The notebook, results, model and every instruction pair are in the lesson's Hugging Face repo.
Ella starts taking requests
Momo is still the model, now tiny-storyteller-v2, which knows Lily, Momo and Zuri, the zebra who joined the class in Lesson 04. Ella is the teacher and the memory. Until now every lesson was "here is a story, guess the next word". That is completion; tonight's goal is following: read a request, then write what it asks for. Ella says a request first, "a story about Zuri being brave in a storm by the watering hole", then reads a story that does exactly that. Request plus story is one instruction pair.
Momo cannot tell where the request ends and the story begins, so Ella holds up two signs (not Lesson 01's three attention cards: two new signs in the vocabulary), one before the request and one when it is his turn. Those are special tokens. And she grades only the story, so during the request her red pen is down: prompt masking.
Where the picture stops being the mechanics: Momo does not understand a request. He learns that words after the first sign, a name or a place, are good words to copy into the story. My check measures that copying and nothing more.
| Term | Momo picture | Real thing |
|---|---|---|
| Instruction tuning | Ella starts every lesson with a request | Fine-tuning on (instruction, response) pairs |
| Instruction pair | A request and the story that answers it | One example: prompt text plus target text |
| Chat template | Ella's fixed order: sign, request, sign, story | A fixed text format that marks who says what |
| Special token | A sign that is not a word | A vocabulary entry the tokenizer never splits |
| Resizing the embedding | Two new rows on Momo's meaning map | Embedding and tied output matrix grow from 4,096 to 4,098 rows |
| Prompt masking | Ella's red pen is down during the request | Request labels set to -100, so the loss skips them |
| Held-out instruction | A request Ella never made in a lesson | A test prompt whose combination never appears in training |
| String match | Did the story say "Zuri" and "by the watering hole"? | Checking that exact substrings appear in the output |
What the notebook does
It loads tiny-storyteller-v2 from Lesson 04's repo, with a local fallback, and builds 4,000 training pairs. Half come from the 27 MB validation file of TinyStoriesInstruct, where each story carries a one-line summary and three words it must use; those become one instruction, "Write a story about this: (summary) Use the words a, b and c." The other 2,000 come from the Momo and Zuri generators, which know what went into each story and write the request themselves: "Write a story about Zuri sharing the shade with Tumi in the long yellow grass." 840 of those 2,000 stories name no place, so their requests name none either; only about 58% of the generator pairs teach place-copying.
Each of the 50 held-out requests, 25 Momo and 25 Zuri, asks for a combination of character, scene and place that appears in no training story. A leakage cell checks that no test or held-out story, no test request and no Lesson 04 quiz story is in the training pairs, and that no training story shows a held-out character-scene-place combination. All five counts were zero.
Every pair becomes one sequence: the end-of-story marker (the token Lesson 01's tokenizer puts after each story), <|instruction|>, the request, <|story|>, the story, the marker. The embedding has one row per token, 4,096 of them, and the new signs are ids 4,096 and 4,097, rows that do not exist yet. So the notebook adds two, starting at the average of the old rows, 5,853,696 parameters in all; because this model's output layer is the same matrix as the embedding (weight tying: Lesson 01's scoring tray reuses the meaning map), it grows too. Labels for every request position are set to -100, the value PyTorch's cross-entropy skips. A widget in the notebook shows one pair's real tokens: 132 of 156 guesses are graded with the mask, all 156 without.
INSTR, STORY = "<|instruction|>", "<|story|>"
tokenizer.add_special_tokens({"additional_special_tokens": [INSTR, STORY]}) # ids 4096 and 4097
model = resize_vocab(model, len(tokenizer)) # 2 new rows; the tied output head grows with them
def encode_pair(instruction, story):
prompt = [eos_id, instr_id] + tokenizer(instruction)["input_ids"] + [story_id]
ids = prompt + tokenizer(story)["input_ids"] + [eos_id]
x = ids[:-1]
y = [t if j + 1 >= len(prompt) else -100 for j, t in enumerate(ids[1:])] # -100: not graded
return x, y
Training is Lessons 01 and 04's loop with batches of 32 padded pairs (short pairs filled to the longest in the batch): 500 steps at a learning rate of 3e-4, each pair seen about four times: 51 seconds on the Mac in the final run, with the GPU shared (37 in an earlier run).
The run and its numbers
| Request set | Both, before | Both, after | Character | Place |
|---|---|---|---|---|
| 50 held-out combinations (the check) | 12% | 44% | 68% → 100% | 14% → 44% |
| 50 familiar combinations | 8% | 86% | 82% → 100% | 8% → 86% |
| 20 cross-world requests | not run | 5% | 100% after | 5% after |
Before tuning, tiny-storyteller-v2 got the request as plain text and kept writing; 6 of 50 stories happened to contain both the character and the place. After tuning, 22 of 50 did: 44%. Every story named the right character, against 68% before; the place was right in 44%, against 14%. The average graded loss on 150 held-out pairs fell from 2.540 to 2.145. On 50 requests whose combination did appear in training the tuned model scored 86%; on 20 that put a character in the other one's world, 1 of 20.
What it means, honestly
The character is the easy half and the place is the hard half. The typical miss has the right character and scene, set in a place the model had practised for that scene: asked for Momo hiding from the rain near the little pond, it hid him in the tall grass. Instruction tuning taught Momo to choose the character and scene from the request, only sometimes to put a new place into a familiar scene, and almost never into the other world.
The check is a string match, and that cuts both ways: a muddled story that says "Zuri" and "by the watering hole" counts, and "near the pond" does not count for "near the little pond". The generator stories are templates, so Momo and Zuri stories are still recited sentences. The real test of general following is a TinyStories-Instruct request: asked for a little dog who finds a lost ball in the park and gives it back to a girl, the model wrote a park, a ball and a girl named Lily, and no dog.
Lesson 04's leftovers show too: before tuning, given a request about Momo hiding from the rain, v2 drifted into a Zuri lesson in which Ella says "Words can hurt, Zuri." No replay was mixed in; measured afterwards, the old quiz went from 11.75 to 12.85.
The frontier uses far more pairs, bigger models and a second stage of human preference training, as in InstructGPT, and instruction tuning across many tasks was shown at scale in FLAN. The mechanics are the same.
Run it yourself
- Colab. Open
lesson_05.ipynbfrom the repo in Google Colab and run all. It loads tiny-storyteller-v2 from Lesson 04's repo; if that is still private, log in with a token that can read it, or run Lesson 04 first and setLOCAL_DIR. On the Mac it took 140 seconds in the final run (102 earlier); it has not been timed on a T4. - Just the model.
tiny-storyteller-instruct/has the weights, the 4,098-token tokenizer andchat_template.json. Generate after<|story|>until the end marker.
What you have now: tiny-storyteller-instruct, 44% of held-out requests right, two special tokens, and a second 23.44 MB Momo beside the v2 you started from. Next, Lesson 06 does the same job by changing 2% of the dials, and asks whether the result is as good.
Check yourself
The number to reproduce: the share of the 50 held-out requests whose story contains both the character and the place, before and after. Mine were 12% and 44%, and held across three local runs; sampling differs across GPUs, so report yours.
- Why is the place so much harder than the character? Look at where places appear in the training stories.
- Train once without the mask. Which number do you expect to change?
- Why was the 86% familiar set not used as the check?
Questions people ask
What is instruction tuning?
Fine-tuning a trained model on request-and-answer pairs in a fixed chat template, so it answers requests instead of just continuing text.
Why mask the prompt out of the loss?
Otherwise the model is also trained to predict the request. Labels of -100 on the request make cross-entropy grade only the answer.
Why do new special tokens need a bigger embedding?
The embedding has one row per token id, and new ids have none. In a weight-tied model the output layer grows with it.
References & Citations
- singhpratech/emc2-lesson-05-instruct on the Hugging Face Hub: lesson_05.ipynb, RESULTS.md,
tiny-storyteller-instruct/, the pairs as JSONL andstory_pairs.py. Run of 11 October 2026. - Starting model: tiny-storyteller-v2 from Lesson 04 (Hub repo
singhpratech/emc2-lesson-04-fine-tune). - Data: roneneldan/TinyStoriesInstruct, validation file (CDLA-Sharing-1.0), from Eldan and Li (2023), TinyStories; plus generated Momo and Zuri stories.
- Wei, J. et al. (2021). Finetuned Language Models Are Zero-Shot Learners. Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback.
- All numbers are from the notebook executed end to end on an Apple Silicon Mac (MPS). The check is a string match on one generated story per request (temperature 0.8, top-k 40, seed = request index).
Subscribe to new posts from theaivibe.org