68,096 × 2,048 × 2 = 278,921,216
278.9M parameters, on representation alone
as a share of the whole model
11%
the other ~2.24B parameters — layers, attention, MLP
why n is 2, and not 1
The vocabulary is charged once on the way in and once on the way out — the input
embedding table and the output projection are separate matrices.
Sarvam's write-up does not say whether they share weights. The shipped config does:
tie_word_embeddings: false
Untied. So two matrices of 139.5M each, not one.