Part 3 trained a 5.8M-parameter GPT from scratch. After 59 seconds of GPU time on the 3.4M-character Holmes train split, it scores 1.626 bits per byte on held-out Holmes text. Gemma 4 E4B (8.0B parameters, pretrained, Apache 2.0) scores 0.953 on the same text with zero training. A QLoRA fine-tune on the Holmes train text takes it to 0.865 in 34 minutes, on the same 8 GB laptop GPU.
Pretraining is the part you cannot afford at home. Fine-tuning is cheap. The most useful lesson in this post is a quantization trap that made my first fine-tune look like a huge win, when most of the “gain” was repairing damage I had done myself. An early plan for this part used Qwen 3.6. It does not fit on this card, and the next section says why.
Train a Tiny GPT at Home series: Part 1: Tokenizers · Part 2: Attention · Part 3: Training · Part 4: Inference · Part 5: Fine-Tuning (you are here)
Full example: Clone the working files at github.com/KingPin/sumguy-examples/tree/main/llm/tiny-gpt/part-5-finetune
Everything ran on the same laptop as Parts 3 and 4: an RTX 3070 Laptop GPU (8 GB), an Intel i7-11800H, and 62 GB of RAM. The software was the pytorch/pytorch:2.14.1-cuda13.2-cudnn9-runtime Docker image with transformers 5.18.0, peft 0.21.2, bitsandbytes 0.50.2, and accelerate 1.15.0. Tested 2026-10-05. The scripts need Part 1’s tokenizer.json and Part 3’s gpu1500.pt.
Why Gemma 4 E4B and Not Qwen 3.6
Qwen 3.6 ships on Hugging Face as Qwen3.6-27B and Qwen3.6-35B-A3B, plus FP8 versions of both. Even at 4 bits per weight, a 27B model needs about 13.5 GB for weights alone, before a single activation. That is not happening on 8 GB, and no amount of cleverness changes the arithmetic.
Gemma 4 E4B is the largest current model I found that fits. Gemma 4 12B exists too, but its 4-bit weights alone take about 6 GB, and E4B already peaks at 7.3 GB in training, so I did not try it. I used google/gemma-4-E4B, the base model, not the -it chat version. We want a model that predicts the next token of raw text, and the instruction-tuned one has been trained to answer questions instead.
Scoring Two Models With Different Alphabets
You cannot compare loss per token here. The tiny GPT has a 4,096-token vocabulary. Gemma has 262,144. One Gemma token covers more text, so its loss per token is a different unit.
Bits per byte fixes that. Take the total loss over a piece of text, divide by the number of bytes in that text, and convert to bits. Lower is better, and the tokenizer drops out of the comparison.
The val text is Part 3’s last 10%, decoded back to text (the BPE roundtrip is exact, so this is the same split Part 3 used). It becomes 578 fixed chunks covering 340,513 scored bytes. Each chunk scores up to 600 characters, with about 200 characters of context in front that the model reads but is not scored on. The body shrinks at whitespace until context plus body fits the tiny model’s 256-token window. Both models see identical context and get scored on identical bytes.
@torch.no_grad()def nats(model, ctx_ids, body_ids, device): """Total -log p of body_ids, given ctx_ids in front of them.""" x = torch.tensor([ctx_ids + body_ids], device=device) out = model(x[:, :-1]) logits = out[0] if isinstance(out, tuple) else out.logits logits = logits[0, len(ctx_ids) - 1 :].float() # the positions that predict the body return F.cross_entropy(logits, x[0, len(ctx_ids) :], reduction="sum").item()
def bits_per_byte(total_nats): n = sum(len(b.encode()) for _, b in val_chunks()) return total_nats / n / math.log(2)The Gemma side tokenizes each chunk’s context with a BOS token, the body without one, and sums nats() over all 578 chunks. Each full val pass takes roughly 100 to 123 seconds.
Squeezing 8 Billion Weights Into 8 GB
The checkpoint is 15 GB on disk. Of its 8.0B weights, 2.8B are the per-layer embedding table (PLE): 262,144 tokens x 42 layers x 256 dims. bitsandbytes quantizes Linear layers. An embedding is not a Linear layer, so the table stays in bf16, which is 5.6 GB by itself. That is most of the card before the real model shows up.
The table is a lookup, though. A forward pass reads a few rows and ignores the rest. So load_gemma() keeps the whole table in CPU RAM and moves only the looked-up rows to the GPU, with two forward hooks.
That took two gotchas to get working.
Gotcha 1: accelerate and the word “cpu”. Put "cpu" in a device_map and accelerate reads it as “store on CPU, execute on GPU”. On every forward pass it tried to copy the whole 5.25 GiB table to the GPU. Out of memory, every time. The fix is to turn accelerate’s dispatch_model into a no-op while the model loads, then restore it and attach the hooks.
Gotcha 2: the checkpoint is multimodal. Gemma4ForCausalLM reported every weight as missing, because the text weights live under model.language_model. in the file. A key_mapping renames them on load, and the vision and audio towers stay on disk.
dispatch, hf_acc.dispatch_model = hf_acc.dispatch_model, lambda *a, **k: Nonetry: model = Gemma4ForCausalLM.from_pretrained( GEMMA, quantization_config=q, dtype=torch.bfloat16, device_map={"model.embed_tokens_per_layer": "cpu", "": 0}, key_mapping={r"^model\.language_model\.": "model."}, )finally: hf_acc.dispatch_model = dispatchtable = model.model.embed_tokens_per_layertable.register_forward_pre_hook(lambda m, args: tuple(a.cpu() for a in args))table.register_forward_hook(lambda m, args, out: out.to("cuda"))The pre-hook sends the token ids to the CPU, where the table lives. The post-hook sends the looked-up rows back to the GPU. With that in place, base-model evaluation peaks at 6,701 MiB in nvidia-smi.
The 4-Bit Trap
My first load_gemma() quantized every Linear layer to NF4, with double quantization and bf16 compute. It loaded, it fit, and it produced these scores on the val text:
base model (all NF4) 1.462 bits per byteLoRA on 10% of text 0.965LoRA on 50% 0.929LoRA on 100% 0.884That looks great. Fine-tuning took a 1.46 model to 0.88, a 40% drop. I almost wrote the post around it.
Then I sampled from the “base” model, and it produced web junk, including “This post was deleted by @…” spam. An 8B pretrained model should not do that, and 1.46 bits per byte is a poor score for one. So I scored the first 40 val chunks again with the base model in bf16 on the CPU, no quantization at all. It got 0.925 bits per byte. The all-NF4 version scored 1.483 on the same 40 chunks.
LoRA had been spending most of its training repairing quantization damage. The “fine-tuning gain” was mostly a fix for my loader.
The cure follows Unsloth’s dynamic 4-bit build of the same model: some layers are too fragile for 4 bits, so they stay in bf16. KEEP_BF16 in common.py lists them.
KEEP_BF16 = [f"model.layers.{i}" for i in (0, 4, 10, 11, 22)] + \ [f"model.layers.{i}.mlp" for i in (1, 5, 6)] + \ [f"model.layers.{i}.self_attn" for i in (2, 3, 5, 6, 9, 23)] + \ ["per_layer_input_gate", "per_layer_projection", "per_layer_model_projection", "lm_head"]That keeps layers 0, 4, 10, 11, and 22 whole (plus 40 and 41, because transformers matches model.layers.4 as a prefix), the MLPs of layers 1, 5, and 6, the attention of layers 2, 3, 5, 6, 9, and 23, and the per-layer projection modules plus lm_head. quant_check.py scores the first 40 chunks three ways:
nf4 everywhere 1.483 bpbnf4 + keep list 1.006 bpbbf16 on CPU 0.925 bpbI also tried keeping only the per-layer projection modules in bf16. That scored 1.156, better than all-NF4 but nowhere near enough. The keep list closes most of the gap to bf16 (1.006 vs 0.925), and the rest is the price of fitting in 8 GB.
Score your quantized base model against an unquantized one before you believe a QLoRA gain. It costs a few minutes and one slow CPU run. Skipping it nearly cost me a wrong article.
The QLoRA Setup
QLoRA freezes the 4-bit base weights and trains small low-rank adapter matrices beside them. Here is the config from finetune.py:
model, tok = load_gemma()model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})model = get_peft_model(model, LoraConfig( r=args.rank, lora_alpha=2 * args.rank, lora_dropout=0.05, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],))Rank 16, alpha 32, dropout 0.05, on all seven projection types. That is 34,881,536 trainable parameters, and the saved adapter is 134 MB on disk. The 8B base stays frozen.
The rest of the recipe:
- Sequences of 512 tokens (BOS plus 511 text tokens), 8 sequences per optimizer step.
- Learning rate 2e-4 with 10% warmup, then cosine decay to 10% of peak. Same shape as Part 3.
- One epoch, no repeats.
- Gradient checkpointing with
use_reentrant=False. - No
prepare_model_for_kbit_training. It casts every weight that is not 4-bit to fp32, which doubles the 5.6 GB table in CPU RAM and every bf16 layer on the GPU.
Training peaks at 7,262 MiB allocated, 7,630 MiB reserved, and 7,653 MiB in nvidia-smi. On an 8 GB card that is tight, so close your browser and anything else holding VRAM. Each optimizer step takes about 10 seconds.
The Scoreboard
I trained both models on 10%, 50%, and 100% of the train text. Part 3’s train.py gained a --frac flag for this, and its default of 1.0 leaves Part 3 unchanged. The smaller tiny runs keep Part 3’s 13.2 passes over their data. Time is training only. Lower bits per byte is better.
| Model | Training | Steps | Time | Val bits per byte |
|---|---|---|---|---|
| Tiny GPT, 10% of train text | from scratch | 150 | 6 s | 2.607 |
| Tiny GPT, 50% | from scratch | 750 | 30 s | 2.344 |
Tiny GPT, 100% (gpu1500.pt) | from scratch | 1,500 | 59 s | 1.626 |
| Gemma 4 E4B base | none | 0 | 0 s | 0.953 |
| Gemma + QLoRA, 10% (77,060 tokens) | QLoRA | 18 | 179 s | 0.925 |
| Gemma + QLoRA, 50% (408,848 tokens) | QLoRA | 100 | 1,005 s | 0.906 |
| Gemma + QLoRA, 100% (826,737 tokens) | QLoRA | 202 | 2,052 s (34 min) | 0.865 |
The 50% tiny run overfit after step 350. train.py keeps the best checkpoint, so the table shows the step-350 score.
Some numbers worth reading slowly:
- Cutting the data from 100% to 10% costs the tiny model 0.98 bits per byte (1.626 to 2.607). It costs Gemma 0.06 (0.865 to 0.925). Gemma already knows English. The data only has to teach it Holmes.
- Gemma’s untrained score beats the fully trained tiny model by 0.67. Pretraining on a huge corpus dwarfs everything you can do with 3.4M characters.
- Fine-tuning on top adds a further 0.088, about 9% (0.953 to 0.865).
- The tiny model trains 35x faster (2,052 s vs 59 s). It also needs about 7x less VRAM: 1,008 MiB allocated at peak, against 7,262 MiB for the QLoRA run.
That is the trade. The small model is cheap and bad. The big one is expensive to run and a lot better before you do anything to it.
Where the Bits Go
I broke the full val set into token types: 87,399 Gemma tokens in all. For the base model, the 68,938 word tokens average 4.31 bits per token and account for 0.873 of the 0.953 total. The 12,636 punctuation tokens average 1.81 bits, and the 5,759 line breaks average 0.72.
After the LoRA fine-tune, words drop to 3.89 bits per token (0.788 bits per byte of share), punctuation to 1.71, and line breaks stay about where they were at 0.74. So almost all of the improvement comes from word choice: which words Conan Doyle picks.
The costliest tokens for both models are ,, the, ., a, and and. Which comma, which “the”: that is where prose gets hard to predict.
Reading the Samples
Scores are one thing. Here is what each model writes. All three used generate.py with temperature 0.8, top_p 0.9, repetition penalty 1.3, 40 tokens, and seed 1. The text is verbatim model output, apart from collapsed whitespace.
Prompt: It was a cold morning in Baker Street whentiny: the case of Merville Hall, for we were very much more than my companion. A broad-man stood as Lestrade entered, and the room came out into his hands on anbase: Sherlock Holmes knocked at the door. "I'm not going to go upstairs," he said, peering through his monocle from under an umbrella as you opened your window for him and let it slipLoRA: Holmes and I started for the scene of our investigation. The weather had been, for days past very stormy indeed; but now there were signs that it would improve before longThe tiny model has the vocabulary of the stories and no plot. The base model writes fluent prose and invents a monocle for Holmes. The fine-tuned one sounds like Watson narrating.
Prompt: Holmes picked up my laptop andtiny: all came back to me. He had a loud brown-eyed, stumpled fish hair man with his thin eyes upon the floor, and he law the little sinbase: said, "I have no idea what you're talking about. I don't know why the old man is so stubborn." "Well… it looks like a bit of trouble has beenLoRA: had a good look at the name, but with his spectacles pulled down he could not make it out. "It is from Mr.-" said I; "I know him as an oldThe laptop does not exist in 1890, and the LoRA model quietly turns it into a document with a name on it. That is a style adapter making the best of an anachronism.
Prompt: The capital of France istiny: the matter, and it does not not know. You will leave me in a madness to be a very disbelieved man, what I do." "How canbase: famous for its fashion and art. This beautiful city with a population numbering over 10 million people boasts countless museums, theatres opera houses – all these attractions are just what you need when planning an educationalLoRA: famous for many things, but the sight which made me first go there was a statue by J. L. Bernat-Guyot—a huge and splendid work in bronze set up onNeither Gemma version said “Paris”. The sculptor is sampled text, so do not cite him. The tiny model has no idea France is a place.
Prompt: To restart a Docker container, youtiny: have no more than yourselves to be married. You will leave me in my knowledge that there is no cause for me to pass yourself." "I should like myself and putbase: can follow these steps: 1. Identify the Container ID 2. Stop and Start Command (Example) docker stop <container_id> # Optional step to ensure that your service isLoRA: can use the <code>docker start</code> command. This will start your stopped docker image and give it access to all of its saved data (that's why I was able to login in myThis last one is the clearest. The LoRA model drags every prompt toward Holmes’s voice, and it still knows docker start. Fine-tuning changed its style, not its knowledge. The tiny model has style fragments and no knowledge at all. These samples use a single seed, so treat them as illustrations only.
Did Gemma Just Memorize Holmes?
Fair worry. Holmes is public-domain Gutenberg text, and big pretrained models have seen plenty of Gutenberg. If Gemma had memorized the val passages, its low score would only measure recall.
I tested it. I took 20 unseen val passages, ran greedy decoding, and checked whether each model reproduces the next 50 characters exactly. The result was 0 of 20 for the tiny GPT, 0 of 20 for the base model, and 0 of 20 for the LoRA model. Gemma has not memorized this val text verbatim. Its lead over the tiny model comes from knowing how English works. Twenty passages is a small sample, so this rules out wholesale memorization and nothing finer.
Running It Yourself
The README uses one docker run pattern for everything. PYTHONUSERBASE and HF_HOME point at folders inside the project, so installed packages and the 15 GB model download survive between containers:
G="docker run --rm --gpus all -v $PWD:/work -w /work/part-5-finetune \ -e PYTHONUSERBASE=/work/pyuser -e HF_HOME=/work/hf \ pytorch/pytorch:2.14.1-cuda13.2-cudnn9-runtime"
$G pip install -q --user --break-system-packages regex==2026.9.29 \ transformers==5.18.0 peft==0.21.2 bitsandbytes==0.50.2 accelerate==1.15.0
$G python eval_bpb.py tiny ../part-3-training/gpu1500.pt$G python eval_bpb.py gemma$G python finetune.py --frac 0.1 --out adapters/frac10$G python finetune.py --frac 1.0 --out adapters/frac100$G python quant_check.py gpuThe --frac 0.1 run takes about 5 minutes. The --frac 1.0 run takes about 37, including the val pass. You need at least 6 GB of free system RAM for the embedding table, and quant_check.py without the gpu argument loads the whole model on the CPU for the bf16 row, which wants about 15 GB for weights alone. Gemma 4 is Apache 2.0 and not gated, so no Hugging Face token is needed.
What This Closes Out
Over five parts, you built a tokenizer, an attention block, a training loop, a KV cache, and a fine-tune. The tiny model is a toy that teaches you where every moving part sits. The fine-tune shows what it takes to get a useful model: borrow someone else’s pretraining, then spend 34 minutes on the part that is yours.
If you want more, two experiments are cheap. Try Gemma 4 E2B, the smaller sibling, with fewer layers in 4-bit, and see how close it gets to its own bf16 score. Or run more than one epoch on the full text and watch where the validation score turns around, as the tiny model’s 50% run did.
Common Questions
Can I fine-tune Gemma 4 on an 8GB GPU?
Yes, the E4B model, with QLoRA. Quantize the Linear layers to 4 bits, keep the 2.8B-parameter embedding table in CPU RAM, and train rank-16 adapters. Peak memory was 7,653 MiB in nvidia-smi on an 8 GB card, so nothing else can share the GPU.
Does QLoRA hurt model quality?
QLoRA can hurt quality badly if every layer goes to 4-bit. Quantizing every layer to NF4 pushed the Gemma base model from 0.925 to 1.483 bits per byte on 40 val chunks. Keeping the most sensitive layers in bf16 brought it back to 1.006. Score your quantized base against an unquantized one first.
How much text do I need to fine-tune an LLM on a writing style?
Less than you would guess. Gemma went from 0.953 to 0.925 bits per byte on 77,060 tokens of Holmes and reached 0.865 at 826,737 tokens. Starting from a pretrained model, 10% of the data got about a third of the total gain. These are one author and one run, so treat them as a rough guide.
Why not fine-tune Qwen 3.6 on 8GB?
Qwen 3.6 ships only as Qwen3.6-27B and Qwen3.6-35B-A3B on Hugging Face, plus FP8 versions. A 27B model needs about 13.5 GB for weights alone at 4 bits, before activations. That exceeds an 8 GB card, so Gemma 4 E4B was the largest current model I found that fit.