Research
We publish the corpus, the evaluation harness and the raw model responses for every study. A score you cannot reproduce is an assertion, not a measurement.
- 2026-09-24Writing India: text-to-image models flatline on Indic scripts
Nine open-weight text-to-image models across Hindi, Bengali, Marathi, Telugu and Kannada. Eight produce a flat 0% exact match on legible Indic signage.
- 2026-09-24What It Actually Said: 16 Indic speech-to-text models
Sixteen ASR models on IndicVoices across 22 languages and 357,000+ utterances. Four produce measurable but meaningless scores by translating instead of transcribing.
- 2026-09-21Bharat Knowledge Probe: what models know about India
19 models on 552 India-specific knowledge items across seven categories. Qwen3.6 27B leads at 53.2% Bharat Score.
- 2026-09-09Can a model advise a farmer? 21 models on Kisan Call Centre questions
500 real crop-advisory questions in 11 Indian languages, graded by an LLM judge on four axes.
- 2026-09-09Indic script reading gap: what vision models miss on Indian signboards
Eight vision-language models read every legible sign in 1,319 unedited photographs from across India.
- 2026-09-09The code-mixing tax: what AI models lose on romanized Hindi
16 models asked the same 1,979 questions five ways. Writing in English letters costs more than mixing languages does.
- 2026-09-09How many tokens does Telugu cost?
Thirteen tokenizers measured across five languages. Dravidian scripts cost four to five tokens per word everywhere.
- 2026-09-09MILU, re-run on the 2026 models
17 models on the complete MILU test set, 79,608 questions across 11 languages with no subsampling, with cost per rupee alongside raw accuracy.
