Skip to content

LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights

Prateek SinghOctober 11, 20269 min read2 views
LoRA Explained by Doing It: Fine-Tune by Changing 2% of the Weights

Lesson 06: LoRA written by hand in a few dozen lines of PyTorch, training 2.11% of a 5.9M-parameter model on the same task as a full fine-tune. The adapter is 0.51 MB instead of 23.44 MB and merges back to within rounding. It was not faster per step, and it followed fewer requests. The real numbers.

Lesson 06 of e=mc² builds LoRA by hand and puts it next to a full fine-tune on the same job. LoRA, low-rank adaptation, is the most common way to fine-tune large models: freeze the model, train a small correction beside each big weight matrix, ship the correction as a file. Here it is a few dozen lines of PyTorch training 2.11% of a 5.9-million-parameter model. The adapter is 0.51 MB instead of 23.44 MB, and a merge (adding the correction into the weights) reproduces every output to within rounding. It was not faster per step, and on this task it followed fewer requests than the full fine-tune.

Where you are: Lesson 05 turned v2 into tiny-storyteller-instruct by moving all 5.85 million dials; 44% of held-out requests came back right, and you now keep two full copies of Momo. Tonight, the same job again touching 2% of the dials, to see what LoRA buys and what it costs.

weights trained
2.11%
123,392 of 5,853,696; rank 5 on every big layer, plus two token rows
file to ship
0.51 MB
the adapter, against 23.44 MB for the full weights
held-out requests, character and place matched
14–26%
LoRA adapters at 3e-3 and 1e-2, against 44% for the full fine-tune
time per step
1–9% less
LoRA against full fine-tuning, four back-to-back timings on a shared GPU; no real speed-up at this size

Run of 11 October 2026 on an Apple Silicon Mac, GPU shared with other jobs. Same pairs, steps, batch and seed as Lesson 05's full fine-tune.

The notebook, four adapters, results and a merge script are in the lesson's Hugging Face repo.

Sticky notes on Momo's dials

Momo is the model, tiny-storyteller-v2 from Lesson 04, 5.85 million dials (the 5.9M of Lessons 04 and 05, plus 512 for the two new token rows). Ella is the teacher and the memory. In Lesson 05 she taught Momo requests by turning all his dials, which left two complete Momos to keep. Tonight she freezes the dials and sticks a small note on each big tray (every large weight matrix; Lesson 01 showed you one, the scoring tray): "when this tray answers, add this correction". Peel the notes off and the old Momo is back; swap in other notes for another skill. That is LoRA, and one set of notes is an adapter.

Each note is two thin strips, A and B. A squeezes what comes in down to a few numbers, the rank r; B spreads them back out. With r = 5, the 256-by-768 attention tray of 196,608 dials gets a note of 5,120 numbers. Alpha sets how loudly the note speaks: the update is scaled by alpha / r = 2.

Where the picture stops being the mechanics: the note is a full-sized matrix built as the product of two thin ones, and it can only change the tray along r directions. LoRA's bet is that a new skill needs only a few directions. This lesson tests the bet.

TermMomo pictureReal thing
LoRA (Low-Rank Adaptation)Sticky notes on frozen dialsA trained low-rank update B·A beside each frozen weight matrix W
AdapterOne set of notes, swappableThe saved A and B matrices, plus any other trained tensors
FreezeNobody may turn these dialsrequires_grad = False: no gradient, no update
Rank rHow many numbers fit through the note's narrow middleInner size of A (r × in) and B (out × r)
AlphaHow loudly the note speaksThe update is scaled by alpha / r
MergeCopy the notes onto the dials, throw the notes awayW := W + (alpha / r) · B · A
The words this lesson adds.

What the notebook does

It loads tiny-storyteller-v2 and Lesson 05's tokenizer, pairs and full fine-tune, then grows the vocabulary by Lesson 05's two special tokens exactly as Lesson 05 did. The heart of it is one class:

class LoRALinear(nn.Module):
    def __init__(self, base: nn.Linear, r: int, alpha: float):
        super().__init__()
        self.base = base
        for p in self.base.parameters():
            p.requires_grad_(False)                                  # the original weights are frozen
        self.scale = alpha / r
        self.A = nn.Parameter(torch.empty(r, base.in_features))
        self.B = nn.Parameter(torch.zeros(base.out_features, r))     # starts at zero: no change at step 0
        nn.init.kaiming_uniform_(self.A, a=math.sqrt(5))

    def forward(self, x):
        return self.base(x) + (x @ self.A.t() @ self.B.t()) * self.scale

    def merged_weight(self):
        return self.base.weight + self.scale * (self.B @ self.A)    # same shape as the frozen weight

B starts at zero, so the wrapped model starts identical to the original; the largest output difference was exactly zero. Each block's four big layers are wrapped. The attention's query-key-value layer is one fused matrix, so it gets one rank-5 adapter shared across Q, K and V, where PEFT, Hugging Face's library for exactly this, gives each its own on models with separate projections. The two new tokens' embedding rows also get a trained correction, 512 numbers, applied to the embedding and the tied output layer alike so it can be merged.

Rank 4 gives 1.69% of 5,853,696 weights, rank 5 gives 2.11%; the notebook uses r = 5 and alpha = 10: 123,392 trainable.

Training matches Lesson 05: same 4,000 pairs, 500 steps of 32, seed, batch order, optimizer and schedule. The learning rate should not match, because a note that starts at zero has to grow, so the notebook trains adapters at 3e-4 (Lesson 05's rate), 1e-3, 3e-3 and 1e-2 and keeps the one with the lowest held-out loss, a rule fixed before the run so the check is never used to pick and stays a test. In full disclosure: the first version tried only the first two, both came out far below the full fine-tune, and the grid was widened to see whether the rate was to blame.

The run and its numbers

ModelTrainableCheckHeld-out loss
Full fine-tune (Lesson 05)5,853,69644%2.145
LoRA r=5, lr 3e-4 (Lesson 05's rate)123,3926%2.274
LoRA r=5, lr 1e-3123,3924%2.229
LoRA r=5, lr 3e-3 (kept: lowest held-out loss)123,39214%2.211
LoRA r=5, lr 1e-2123,39226%2.216
The check: share of 50 held-out requests whose story contains the named character and the exact place (string match), the same requests and seeds as Lesson 05. Held-out loss: average graded loss on 150 held-out pairs.
CostFull fine-tuneLoRA r=5
Time per training step (back to back)104.7 ms95.6 ms
Saved file23.44 MB0.51 MB
AdamW state (computed: 2 × 4 bytes per trainable weight)46.8 MB0.99 MB
Timing: 100 steps of each, same batches, after 10 warm-up steps, back to back on a shared GPU, final run (earlier runs: 137.0 / 133.5, 183.8 / 181.9 and 64.3 / 62.7 ms). Optimizer state is arithmetic, not a measurement.

At Lesson 05's learning rate the adapter learned almost nothing of the request-following: 6% of held-out requests, no better than the untuned model's 12%. Higher rates helped, 14% at 3e-3 and 26% at 1e-2, against 44% for the full fine-tune. The held-out loss barely separated them, 2.211 against 2.216, so the adapter kept by the rule, the 3e-3 one, is not the one with the best check. Every check and loss number was identical in the last three runs of the notebook.

The merge worked as advertised: folding each B·A into its frozen weight gave a plain model of the original shape whose outputs differed from the adapter model's by at most 3.29e-05, and all 50 generated test stories were identical. The standalone merge_lora.py rebuilt the same weights to within 1.49e-08.

What it means, honestly

LoRA did not match full fine-tuning here. Following a request means reading it and choosing from it, and a rank-5 correction learned that choice less well than 5.85 million free weights did in the same 500 steps. LoRA is usually reported close to a full fine-tune on much larger models; I did not test that here.

The speed result is the one people most often get wrong. LoRA skips the weight-gradient matrix multiplications of the frozen layers and their optimizer step, but the forward pass and the backward pass through the activations, the numbers each layer produces on the way up, still run through every layer, so even at scale the per-step saving is real but modest. Here LoRA steps were 1% to 9% faster across four timings (95.6 against 104.7 milliseconds in the final run), with other jobs sharing the GPU. What LoRA saves is memory and storage: optimizer state of 0.99 MB against 46.8 MB, and a file 2.2% of the size. On billion-parameter models those savings decide whether fine-tuning fits on one GPU, which is why the original LoRA paper made them its headline.

Run it yourself

  • Colab. Open lesson_06.ipynb from the repo in Google Colab and run all. It needs Lesson 04's and Lesson 05's Hub repos; if they are still private, use a token that can read them, or run those lessons first. On the Mac, GPU shared, it took 11 minutes in the final run (6 to 8 in the earlier runs); it has not been timed on a T4. On a free T4 this may run close to the 10-minute mark; if it does, drop one learning rate from the grid.
  • Just the adapter. python merge_lora.py --adapter adapters/lora_r5_lr0.003.pt --out merged (merge_lora.py) gives a plain model; use Lesson 05's tokenizer and template.

What you have now: a 0.51 MB LoRA adapter for the Lesson 05 job that merges back to within rounding, and the honest number: 14 to 26% against 44%. Stage B is done. Next, Stage C: Lesson 07 gives a new, tiny Momo a real job inside pankhllm, choosing operations instead of writing stories.

Check yourself

The numbers to reproduce: 123,392 trainable weights (2.11%), an adapter near 0.51 MB, step times within about 10% of full fine-tuning, and a check well below the full fine-tune's. The count and size are exact; the check moves by a few points between runs.

  1. Why does B start at zero and A at random?
  2. Why would LoRA save some time per step on a 7-billion-parameter model?
  3. Held-out loss and the check picked different adapters. Which should choose, and why?

Questions people ask

Is LoRA faster than full fine-tuning?

Per step, barely: it skips frozen weights' gradients and updates, not the passes through them. Here it was 1% to 9% faster across four timings. It saves optimizer memory and file size.

Does LoRA match full fine-tuning?

Not here: 44% for the full fine-tune, 26% for the best adapter on the check (lr 1e-2), which is not the kept one, chosen by held-out loss.

What does merging do?

It adds the scaled B·A into each frozen weight, giving an ordinary model of the original shape. Merged and unmerged models wrote identical stories for all 50 requests.

References & Citations

Subscribe to new posts from theaivibe.org

No spam — just new posts. One-click unsubscribe.
Share this article