~/glossary
Glossary: benchmark terms and metrics
Every metric and term used across our benchmarks, defined in full, with the method behind it.
- Bharat Score
The macro-mean accuracy a language model achieves across the seven India-specific categories of BKP-500.
- Code-mixing tax
The accuracy a model loses when a question arrives in mixed or romanized script rather than a single native script.
- Tokenizer fertility
The average number of tokens a tokenizer produces per word of input text, and so the price of every request.
For how the corpora behind these metrics are built and graded, see our methodology.
