Both units shown. Fertility is tokens per word, so it depends on what counts as a
word. Bytes per token needs no such notion — higher means each integer carries more text.
LANGUAGEWORDS
SARVAM-1
o200k · GPT-4o
cl100k · GPT-4
FERTB/TOK
FERTB/TOK
FERTB/TOK
English4,000
1.693.68
1.354.60
1.374.53
Hindi4,000
1.658.52
1.877.50
5.342.62
Bengali4,000
2.078.81
2.537.21
8.422.17
Tamil3,865
2.539.47
3.187.52
11.782.03
Telugu4,000
2.507.89
3.146.28
12.561.57
Kannada4,000
2.589.20
3.486.81
15.111.57
Indic avg
2.268.78
2.847.06
10.641.99
WHAT THE VENDOR NUMBER LOOKS LIKE HERE
Sarvam publishes a fertility range of 1.4 to 2.1. On this corpus it measures
2.26 — just outside it. A single matched paragraph had put it at 1.40, the
exact bottom of the range. Fertility is corpus-dependent, and that is the finding: a number
like this means nothing without the text it was measured on.
THE TRADE, VISIBLE IN ONE ROW
On Indic text Sarvam packs 8.78 bytes into every token against
o200k's 7.06. On English it manages only 3.68 against o200k's
4.60. Same tokenizer, better on one and worse on the other —
that is what spending a vocabulary budget looks like from both sides.