A language model never sees your words. Before the first matrix multiplication, your text becomes a list of integers, and those integers become vectors. Everything the model knows is built on top of those two conversions.
You can build the first half in about 100 lines of Python. The result holds its own against OpenAI’s tokenizers on the text it was trained on, and loses badly on everything else. That gap is the lesson of this post: a tokenizer encodes the assumptions of its training data. This is Part 1 of a series where we build and train a small GPT-style model from scratch on home hardware (a CPU and an 8GB consumer GPU), with measured numbers. Part 2 covers attention and positional encoding. Part 3 builds the full model and the training loop, CPU against GPU. Part 4 covers inference: KV cache and sampling.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/…/tiny-gpt/part-1-tokenizer
Why Your Model Can’t Just Read
A neural network does arithmetic on numbers. Text is not numbers, so something has to bridge the gap. You have three obvious choices, and two of them are bad.
Characters. Give every letter an id. The vocabulary is tiny, and nothing is ever unknown. But “unquestionably” is 14 steps in the sequence, and a model with a fixed context window spends it all on spelling. Attention cost grows with sequence length, so long sequences hurt twice.
Whole words. Give every word an id. Sequences are short, but the vocabulary explodes (“telegram”, “telegrams”, “telegraphed” are three rows). Any word you did not see in training becomes <unk>, and the model cannot read it at all. Typos, names and code identifiers all break it.
Subwords. Pick pieces in between. Common words get one id. Rare words get split into familiar chunks. And if you start from raw bytes as the floor, every possible string can be encoded, because every string is bytes. Nothing is ever unknown.
Byte-pair encoding (BPE) is the standard way to get there. Start with 256 ids, one per byte. Find the most frequent adjacent pair in your training text. Glue it into a new id. Repeat until you hit the vocabulary size you want. GPT-2, GPT-4 and GPT-4o all use a version of this.
The Corpus: Nine Detective Novels
The series trains on one small, fixed corpus: nine public-domain Sherlock Holmes books from Project Gutenberg. It comes to 3,793,263 bytes. get_data.py downloads and cleans it.
Why so small? A fixed corpus makes every number in this series reproducible. Run the scripts and you get my numbers.
It also makes the tokenizer’s bias easy to see. The training text is Victorian English prose. Hold that thought.
How BPE Training Works
Four steps: split with a regex, convert to bytes, count pairs, merge the winner. Here is the split pattern from bpe.py:
import regex
SPLIT = regex.compile( r"""'(?:[sdmt]|ll|ve|re)| ?\p{L}+| ?\p{N}+| ?[^\s\p{L}\p{N}]+|\s+(?!\S)|\s+""")Note the import. This is the third-party regex package, not the stdlib re. The pattern uses \p{L} (any Unicode letter) and \p{N} (any Unicode digit), and the stdlib module does not support those classes. Install it with pip and move on.
The pattern cuts text into chunks: contractions, letter runs, digit runs, punctuation runs and whitespace. A leading space stays glued to the word after it, so you get " Holmes", not " " plus "Holmes". That matters. Merges never cross a chunk boundary, so the tokenizer can never learn a token that spans two words.
Now the training loop:
@classmethoddef train(cls, text: str, vocab_size: int, verbose: bool = False) -> "BPE": # Deduplicate chunks first: " the" appears ~32k times but only needs # to be stored once with a count. This is what keeps training fast. words = Counter(tuple(c.encode("utf-8")) for c in SPLIT.findall(text)) tok = cls() for new_id in range(256, vocab_size): counts = pair_counts(words) if not counts: break best = max(counts, key=counts.get) words = {merge(w, best, new_id): f for w, f in words.items()} tok.merges.append(best) tok.ranks[best] = new_id tok.vocab[new_id] = tok.vocab[best[0]] + tok.vocab[best[1]] ... return tokThe Counter on the first line is the trick that makes pure Python viable. Most chunks repeat constantly. Store each distinct chunk once with a count, and every pair count gets multiplied by that frequency. You scan the unique words each round, not the whole book.
Encoding replays the merges in the order they were learned:
def _encode_chunk(self, chunk: str) -> list[int]: if chunk in self.cache: return self.cache[chunk] ids = tuple(chunk.encode("utf-8")) while len(ids) > 1: # Apply the earliest-learned merge present in this chunk. pair = min(zip(ids, ids[1:]), key=lambda p: self.ranks.get(p, 1 << 30)) if pair not in self.ranks: break ids = merge(ids, pair, self.ranks[pair]) self.cache[chunk] = list(ids) return self.cache[chunk]Take the pair with the lowest rank (learned earliest), merge it, and repeat until no known pair is left. The per-chunk cache means " the" is only merged once, then looked up for the other 32,150 occurrences.
What It Learned
The training script asks for 4,096 tokens: 256 bytes plus 3,840 merges. Here are the first merges, verbatim:
merge 256: b' ' seen 104,560xmerge 257: b' t' seen 84,595xmerge 258: b'he' seen 72,050xmerge 259: b' a' seen 62,423xmerge 260: b'in' seen 46,330xmerge 261: b' w' seen 44,537xmerge 262: b' s' seen 39,518xmerge 263: b' the' seen 39,084xThe first merge is two spaces, 104,560 times. That is Project Gutenberg’s indented text: the tokenizer learned the file format before it learned English. Then come the usual suspects (' t', 'he', ' a', 'in'), and the word ' the' finally forms at merge 263, after ' t' and 'he' had to exist first.
Later merges get weirder as the counts drop:
merge 500: b'han' seen 1,582xmerge 1000: b'ssible' seen 361xmerge 1500: b' sound' seen 178xmerge 2000: b'pro' seen 114xmerge 2500: b'retched' seen 79xmerge 3000: b' admir' seen 59xmerge 3500: b' occup' seen 46xmerge 4000: b'two' seen 37xBy merge 4000, a pair that shows up only 37 times in 3.8 MB earns a slot. Diminishing returns are real, and this is where you see them. Past some point, you are memorizing rare strings.
The run took 140.5 seconds on one CPU core. That is pure Python, single-threaded, no tricks beyond the dedupe. The finished corpus encodes to 1,030,773 tokens, which is 3.68 bytes per token.
Now a sentence it has never seen: Holmes examined the telegram. It was unquestionably from Moriarty.
14 tokens: Holmes| examined| the| telegram|.| It| was| un|quest|ion|ably| from| Moriarty|." Moriarty" is a single token. The corpus knows its villain. “unquestionably” is not a common enough word to earn a slot, so it falls apart into four pieces: un, quest, ion, ably. Still decodable, still lossless, just a little more expensive.
The longest tokens in the vocabulary read like the table of contents of a Victorian detective novel: sixteen spaces, extraordinary, investigation, circumstances, advertisement, disappearance, considerable, conversation, professional, astonishment. Nobody designed that vocabulary. It fell out of counting.
Our 4k Tokenizer vs OpenAI’s
compare.py runs the same snippets through our tokenizer and three from the tiktoken library. The mapping, so the names mean something: gpt2 is the GPT-2 tokenizer, cl100k_base is GPT-4 and GPT-3.5-turbo, and o200k_base is the GPT-4o family.
Token counts per snippet (fewer is better):
| snippet | ours-4k | gpt2 | cl100k_base | o200k_base |
|---|---|---|---|---|
| holmes prose | 31 | 31 | 31 | 31 |
| tech prose | 41 | 26 | 25 | 25 |
| python | 61 | 82 | 36 | 37 |
| compose yaml | 57 | 59 | 41 | 41 |
| german + emoji | 44 | 29 | 20 | 17 |
| vocab size | 4,096 | 50,257 | 100,277 | 200,019 |
On Holmes prose, a 4,096-entry vocabulary ties a 200,019-entry one. Same text, same 31 tokens. That is what training on your own data buys you.
Step off the Victorian road and it falls apart. Tech prose costs 41 tokens against 25. German with an emoji costs 44 against 17, because our corpus has almost no German and no emoji, so those strings decompose to raw bytes.
The Python snippet is the interesting one. Look at the first line:
ours-4k de|f| ret|ry|(|f|n|,| attempt|s|=|3|)|:gpt2 def| ret|ry|(|fn|,| attempts|=|3|):cl100k_base def| retry|(fn|,| attempts|=|3|):o200k_base def| retry|(fn|,| attempts|=|3|):Our tokenizer never saw def in a Holmes novel, so it spells it out. gpt2 has def but needs two tokens for retry. The newer tokenizers do both in one.
gpt2 is also the worst on Python overall (82 tokens, worse than ours at 61). The reason is whitespace. gpt2 has no merges for runs of spaces, so indentation costs one token per space. Our tokenizer learned the double space from Gutenberg’s indented paragraphs, so it gets a small break. OpenAI fixed this in later tokenizers. In this snippet, 45 of gpt2’s 82 tokens are pure whitespace, against 5 of cl100k_base’s 36. That fix accounts for most of the drop.
Now the full corpus, 3,793,263 bytes of Holmes:
ours-4k 1,030,773 tokens 3.68 bytes/token 0.59sgpt2 1,090,642 tokens 3.48 bytes/token 0.22scl100k_base 896,678 tokens 4.23 bytes/token 0.24so200k_base 892,725 tokens 4.25 bytes/token 0.15sOur 4k vocabulary beats gpt2’s 50k vocabulary on the text it trained on. It loses to cl100k_base and o200k_base by about 15 percent, with a vocabulary 24 to 49 times smaller. On the wrong text, it loses by a lot more.
Speed is the other column. tiktoken is Rust, so it is faster by a factor of two to four even against our cache-backed Python loop. For a 3.8 MB corpus that is a rounding error.
Why did the big labs keep growing vocabularies? Fewer tokens per document means more text fits in a context window, and API pricing is per token. A tokenizer that packs 4.25 bytes per token instead of 3.48 gives you about 22 percent more room, for free. If you want to see how that interacts with limits, Context Window vs Token Limit covers the accounting.
From Integers to Vectors: Embeddings
The second half of the pipeline. A token id is just a label. The model needs a vector it can do math on, and it gets one from an embedding table: a matrix with one row per token, d_model columns each.
An embedding lookup is row indexing, nothing fancier. embeddings.py proves it two ways:
ids = torch.tensor(tok.encode("Holmes examined the telegram."))vectors = emb(ids)print(f"{len(ids)} token ids {ids.tolist()} -> tensor of shape {tuple(vectors.shape)}")assert torch.equal(vectors, emb.weight[ids])one_hot = F.one_hot(ids, VOCAB).float()assert torch.allclose(vectors, one_hot @ emb.weight)5 token ids [1207, 1997, 263, 2307, 46] -> tensor of shape (5, 256)lookup matches weight[ids] and one_hot @ weightFive ids in, a 5 by 256 matrix out. The one-hot multiply is the textbook definition, and indexing is the shortcut that skips multiplying by a pile of zeros. Both give identical results, and the script asserts it.
Before training, those rows are random. Here is cosine similarity between a few tokens in the untrained table:
cosine(' Holmes', ' Watson') = +0.037cosine(' Holmes', ' telegram') = +0.035cosine(' Watson', ' telegram') = -0.019Holmes is no closer to Watson than to telegram. Both are noise around zero. Meaning has to be learned, and it only appears once the model trains. Part 3 reruns this exact check after training, so you can see the numbers move.
What the Table Costs
The embedding table has vocab_size * d_model parameters. This is where tokenizer choice stops being about token counts and starts being about memory:
| vocab | d_model | params | fp32 MB |
|---|---|---|---|
| 4,096 | 256 | 1,048,576 | 4.2 |
| 4,096 | 768 | 3,145,728 | 12.6 |
| 50,257 | 256 | 12,865,792 | 51.5 |
| 50,257 | 768 | 38,597,376 | 154.4 |
| 100,277 | 768 | 77,012,736 | 308.1 |
| 200,019 | 768 | 153,614,592 | 614.5 |
A 200k vocabulary at d_model 768 is 153 million parameters in the embedding table alone. That is several times bigger than the whole model this series trains in Part 3, which stays under 30 million parameters. A tiny home-trained model cannot afford a giant vocabulary: most of its parameters would be a lookup table, and most rows would get too few training examples to learn anything.
That is why we train a 4k vocabulary on the model’s own corpus. The tokenizer costs 4.2 MB at d_model 256, and every row sees plenty of examples.
Run It Yourself
You do not need a GPU for this part. I tested it in October 2026 with torch 2.14.1 (CPU build), tiktoken 0.14.0, and Python 3.13 and 3.14. Clone the examples repo, then run the series setup from llm/tiny-gpt/:
python3 -m venv .venv && source .venv/bin/activatepip install torch==2.14.1 --index-url https://download.pytorch.org/whl/cpupip install -r requirements.txtpython get_data.py # writes data/holmes.txtThen, from part-1-tokenizer/, run the four scripts in order:
cd part-1-tokenizerpython bpe.py # self-check: round-trips, unseen text still workspython train_tokenizer.py # ~140s on one core, writes tokenizer.jsonpython compare.py # token counts vs tiktokenpython embeddings.py # lookup, similarity, cost tablebpe.py ends with an assert-based self-check. It trains on a toy string, round-trips three inputs including an emoji and an empty string, and checks that merges shrink the sequence. If that passes, the tokenizer is sound.
What to Take Away
A tokenizer is a compression scheme fitted to a corpus. It looks neutral, and it is not. Ours learned that double spaces are common and that Moriarty is a person, and it learned nothing about def, YAML or German. The big OpenAI tokenizers have their own biases: they were fitted to web text, code and multiple languages, and you can read the priorities in the token counts.
For a tiny model, this settles the design. Use a small vocabulary trained on your own data, because the embedding table is the one cost you control. For an API model you do not train, remember that token counts differ by content: the same token budget holds less German than English.
Next up is Part 2, attention. The embeddings from this post go in, and the model finally gets a way to look at other positions in the sequence. We will build it in PyTorch and measure what it costs.
Common Questions
How many tokens is a word in GPT-4?
Most common English words are one token in GPT-4’s cl100k_base tokenizer, leading space included. On the Holmes corpus, cl100k_base averages 4.23 bytes per token, about one short word. Rare words and names split into several tokens, and code and non-English text cost more tokens per word than English prose.
Should I train my own tokenizer or use tiktoken?
Use tiktoken for anything that talks to an OpenAI model, because token counts must match the API. Train your own only when you train a model from scratch on a narrow corpus. A small custom vocabulary shrinks the embedding table and fits your domain, as the 4k run here shows.
Why does a bigger vocabulary use fewer tokens?
A bigger vocabulary has more multi-character tokens, so each token covers more bytes. On the Holmes corpus, o200k_base reaches 4.25 bytes per token and our 4,096-token vocabulary reaches 3.68. The tradeoff is table size: 200,019 tokens at d_model 768 is 153,614,592 parameters.
Do I need a GPU to train a BPE tokenizer?
No. BPE training is counting and merging, not matrix math, so a GPU does not help. The tokenizer here trained 4,096 tokens in 140.5 seconds on one CPU core in single-threaded Python. Libraries written in Rust train far larger vocabularies on far larger corpora, also on CPU.