Fertility and compression on a real corpus — 4,000 words of Wikipedia prose per language, not one matched paragraph
Both units shown. Fertility is tokens per word, so it depends on what counts as a word. Bytes per token needs no such notion — higher means each integer carries more text.
LANGUAGEWORDS SARVAM-1 o200k · GPT-4o cl100k · GPT-4
FERTB/TOK FERTB/TOK FERTB/TOK
English4,000 1.693.68 1.354.60 1.374.53
Hindi4,000 1.658.52 1.877.50 5.342.62
Bengali4,000 2.078.81 2.537.21 8.422.17
Tamil3,865 2.539.47 3.187.52 11.782.03
Telugu4,000 2.507.89 3.146.28 12.561.57
Kannada4,000 2.589.20 3.486.81 15.111.57
Indic avg 2.268.78 2.847.06 10.641.99
WHAT THE VENDOR NUMBER LOOKS LIKE HERE
Sarvam publishes a fertility range of 1.4 to 2.1. On this corpus it measures 2.26 — just outside it. A single matched paragraph had put it at 1.40, the exact bottom of the range. Fertility is corpus-dependent, and that is the finding: a number like this means nothing without the text it was measured on.
THE TRADE, VISIBLE IN ONE ROW
On Indic text Sarvam packs 8.78 bytes into every token against o200k's 7.06. On English it manages only 3.68 against o200k's 4.60. Same tokenizer, better on one and worse on the other — that is what spending a vocabulary budget looks like from both sides.
Measured with corpus_experiment.py, published with this article·@sushant_p18