TOKENIZATION
/ ONE SENTENCE, FOUR SCHEMES
COMPRESSION · BYTES / TOKEN
Every scheme maps the
same 11 bytes
to integers.
The question is only:
how many, and how robust?
"the cat sat"
Character
Unicode code points
t
h
e
␣
c
a
t
␣
s
a
t
1.0
B/tok
~150K vocab · mostly unused
Byte
raw UTF-8
74
68
65
20
63
61
74
20
73
61
74
1.0
B/tok
256 vocab · sequences explode
Word
split on whitespace
the
cat
sat
→ but "catsat", typos, names →
UNK
~5
B/tok
✕ ∞ vocab · UNK unrecoverable
THE NEGOTIATED PEACE
BPE
learned merges · GPT-2
the
▁
cat
▁
sat
leading space is part of the token
~4
B/tok
✓ 50,257 vocab · zero UNK
FIG. 1 — CHAR & BYTE COMPRESS BADLY · WORD BREAKS ON THE UNKNOWN ·
BPE KEEPS BOTH
SOURCE: STANFORD CS336 L1 · DIAGRAM: @SUSHANT_P18