~/benchmarks
Benchmarks
corpora we build, maintain and re-run
These are evaluation suites we built and maintain, each with its own corpus, grading method and published metrics. They are distinct from our research studies, which measure models against existing or one-off datasets. Each page here explains what the benchmark tests and how to get the data.
- BKP-500: the Bharat Knowledge Probe
552 India-specific items across seven categories: lakh and crore arithmetic, informal weights, state land units, crop calendars, fiscal-year conventions, government schemes and structural identifiers. Graded deterministically by code. Latest run: 19 models, 2026.
- Indic-KCC: the Indian crop-advisory benchmark
500 real Kisan Call Centre questions presented in 11 Indian languages, graded on correctness, naturalness, groundedness and safety. Latest run: 21 models, 2026.
