TOKENIZATION / ONE SENTENCE, FOUR SCHEMES
COMPRESSION · BYTES / TOKEN
Every scheme maps the same 11 bytes to integers.
The question is only: how many, and how robust?
"the cat sat"
CharacterUnicode code points
the cat sat
1.0 B/tok
~150K vocab · mostly unused
Byteraw UTF-8
746865 20636174 20736174
1.0 B/tok
256 vocab · sequences explode
Wordsplit on whitespace
thecatsat → but "catsat", typos, names → UNK
~5 B/tok
✕ ∞ vocab · UNK unrecoverable
THE NEGOTIATED PEACE
BPElearned merges · GPT-2
the cat sat leading space is part of the token
~4 B/tok
✓ 50,257 vocab · zero UNK
FIG. 1 — CHAR & BYTE COMPRESS BADLY · WORD BREAKS ON THE UNKNOWN · BPE KEEPS BOTH
SOURCE: STANFORD CS336 L1 · DIAGRAM: @SUSHANT_P18