We ran 21 models against BKP-500 — 552 India-specific core items across 7 categories (numerals, weights/volumes, land units, agricultural seasons, fiscal-year conventions, government schemes, structural identifiers), each paired 1:1 against an international control twin (532 control items), under two prompt regimes and 3 repeated samples per item. Qwen3.6 27B leads the roster at 53.2% Bharat Score; Sarvam-M 24B is the strongest India-built/tuned model at 45.1%. The single weakest spot for the roster as a whole is Land units by state — the worst-scoring category for 15 of 21 models.
BKP-500 tests whether a model actually knows India-specific conventions — lakh/crore numerals, informal weights (tola, seer, chhatak), land units that vary by state, kharif/rabi/zaid crop calendars, the April–March fiscal year, and government-scheme facts — rather than defaulting to the international convention that dominates most pretraining data. Every core item is paired with a matched control item that asks the equivalent question under an international/self-contained convention, so a model's Locale Gap Delta (control accuracy minus India accuracy) isolates a genuine India-specific knowledge gap from general capability.
Qwen3.6 27B leads the roster at 53.2% Bharat Score (95% CI 48.9–57.3%), a global open-weight model, not an India-built one. Sarvam-M 24B is the strongest India-built/tuned model at 45.1%, landing at rank 3 overall. Across the full roster, global open-weight models average 35.2% against India-built models' 20.6% — a 14.6-point gap, though with only 10 and 11 models per bucket that comparison's own confidence interval is wide (see the bucket comparison below). Land units by state is the hardest category for the roster as a whole: the single lowest-scoring category for 15 of 21 models, well ahead of any other category.
Qwen3.6 27B's lead over gpt-oss-20b (6.3 points of Bharat Score) and gpt-oss-20b's lead over Sarvam-M 24B (1.7 points) both survive Holm-Bonferroni correction across every adjacent-rank test at once (0.05 family-wise alpha). From rank 3 (Sarvam-M 24B) through rank 10 (Gemma 3 12B), no adjacent pair is statistically distinguishable — a genuine 8-model plateau. It doesn't hold for the rest of tier 1, though: 3 more Holm-significant breaks punctuate the remaining ranks (Gemma 3 12B→Llama 3.1 8B Instruct, Qwen2.5 7B→Gemma 3 4B, Gemma 3 4B→Navarasa 2.0 7B) before the tier 1/tier 2 cliff itself (see below).
Sarvam-M 24B scores 45.1%, tier 1 — 1.7 points behind #2 gpt-oss-20b (a Holm-significant gap), 3.3 points ahead of #4 Gemma 3 27B (not a Holm-significant gap). It shares tier 1's upper half with 1 other India-built/tuned model (Krutrim-2 12B).
Land units by state is the worst-scoring category for 15 of 21 models — the next most common weak spot is a distant runner-up. State-by-state variation in land-unit definitions (a bigha in Punjab isn't a bigha in Bihar) appears to be genuinely harder for every model family we tested than any other category in the corpus.
Navarasa 2.0 7B (20.7%, tier 1) and Krutrim-1 7B (14.4%, tier 2) sit only 6.3 points apart in raw score, but that gap is Holm-significant (p<0.001) — tier 2 isn't an artifact of the tiering method, it's a real capability floor 7 of 21 models fall below.
Gemma 3 12B INT4 shows the largest control-favoring gap in the roster by far (+29.6 points, next-highest is under 4) — but it's also the roster's lowest-scoring model overall (4.6%, rank 21 of 21), so this extreme reading is more likely a symptom of that model's broader breakdown than a clean locale-specific signal. Qwen3.6 27B is a cleaner example of the gap moving independently of accuracy: it's rank 1 of 21 by Bharat Score (53.2%), yet it's actually 14.0 points easier on India-framed items than on the matched international control — the reverse of the expected direction, from a model that's clearly not broken.
Consistency — agreement across a model's own repeated samples — sits at 87.5–100.0% across the entire roster, regardless of Bharat Score. A model scoring 5% and one scoring 50% are both answering the same way every time they're asked — low accuracy here means confidently, repeatably wrong, not random guessing.
Gemma 3 12B scores 35.1% (rank 10 of 21). The same model, quantized to INT4, scores 4.6% — dead last, rank 21 of 21, below even Gemma 3 4B (25.3%). Its refusal rate jumps to 46.5% and it scores a flat 0.0% on 5 of the 7 categories — this reads as a broken quantization, not a graceful capability trade-off.
Qwen3.6 27B loses 20.7 points going from natural-language (R2) to strict-JSON (R1) answers, Sarvam-M 24B loses 17.0, gpt-oss-20b loses 14.0 — and these are exactly the roster's top 3 ranks by Bharat Score. Meanwhile Airavata 7B actually gains 7.8 points under strict JSON. Format compliance and raw capability pull in opposite directions here: the strongest models seem to lose more to answer-extraction friction, not less.
Sarvam 30B leads the roster on Ambiguity Handling (69.4%) with a low Overconfidence rate (8.3%), despite ranking 9 of 21 overall. At the other extreme, Param-1 7B and Gemma 3 12B INT4 both score 0% Ambiguity Handling — never hedging on a deliberately underspecified item — but for very different reasons: Param-1 7B answers confidently and wrongly (58.3% overconfidence), while Gemma 3 12B INT4's own 0.0% overconfidence rate suggests it's mostly just not answering at all.
Sarvam 30B, Sarvam-M 24B and Krutrim-1 7B post the roster's top 3 Unit Discipline scores (75.4%, 74.7%, 71.1%) — all three India-built/tuned, despite mixed overall ranks (9, 3, 15 of 21). Airavata 7B and Param-1 7B trail furthest behind (14.5%, 19.4%) — both are also among the roster's highest-refusal models, so a chunk of their missing units is likely missing answers, not wrong ones.
Qwen2.5 7B (30.9%, rank 12) → Qwen3-VL 8B (39.2%, rank 5, +8.3pp at near-identical size) is a generational jump, not a scale one. Qwen3.6 27B (53.2%, rank 1) then adds real scale on top, landing as the roster's outright leader. Every step in this family's chain moved in the same direction — no other model family in the roster shows as clean a progression.
Global open-weight models lead India-built/tuned models by 23.5 points on Indian numeral system and 21.7 on Weights, volumes & informal measures — the two categories built on precise lakh/crore-scale arithmetic. The gap narrows to 8.6 points on Agricultural seasons & crop calendars, the smallest of any category. The India-built/tuned bucket doesn't trail evenly across categories: it closes ground fastest on categories that reward cultural/domain familiarity over precise numeric conversion, and falls furthest behind exactly where scale arithmetic dominates.
Ranked by Bharat Score (macro-mean accuracy across all 7 categories), tiered by chained bootstrap-CI overlap. Locale Gap Delta = control-item accuracy minus India-item accuracy on matched pairs; positive means the international framing was easier for that model, negative means the India-specific framing was.
| tier | model | bharat score (95% CI) | locale gap Δ | oom error rate | unit discipline | refusal rate | consistency |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.6 27B global | 53.2% (48.9-57.3) | -14.0% | 23.0% | 69.8% | 0.7% | 100.0% |
| 1 | gpt-oss-20b global | 46.8% (43.0-50.9) | -9.4% | 25.8% | 65.7% | 0.1% | 89.3% |
| 1 | Sarvam-M 24B india | 45.1% (40.9-49.1) | -8.6% | 32.8% | 74.7% | 3.6% | 92.6% |
| 1 | Gemma 3 27B global | 41.8% (38.3-45.8) | -5.8% | 27.8% | 68.6% | 0.1% | 99.6% |
| 1 | Qwen3-VL 8B global | 39.2% (35.1-43.4) | -1.8% | 28.7% | 66.9% | 0.0% | 97.6% |
| 1 | Mistral Small 3.1 24B global | 38.8% (35.0-42.9) | -3.5% | 26.7% | 61.0% | 0.4% | 98.2% |
| 1 | Krutrim-2 12B india | 38.7% (34.6-43.3) | -5.8% | 32.5% | 66.4% | 0.2% | 99.6% |
| 1 | Phi-4 14B global | 38.7% (34.8-43.0) | -4.0% | 29.2% | 68.1% | 3.7% | 99.0% |
| 1 | Sarvam 30B india | 37.1% (33.5-40.9) | -6.3% | 40.0% | 75.4% | 2.3% | 87.5% |
| 1 | Gemma 3 12B global | 35.1% (31.3-39.4) | -2.8% | 32.4% | 66.7% | 0.0% | 99.7% |
| 1 | Llama 3.1 8B Instruct global | 32.5% (28.5-36.6) | -3.1% | 40.4% | 63.5% | 2.2% | 100.0% |
| 1 | Qwen2.5 7B global | 30.9% (27.1-34.8) | +1.2% | 31.6% | 67.2% | 0.1% | 98.1% |
| 1 | Gemma 3 4B global | 25.3% (22.1-28.8) | -2.0% | 41.3% | 60.9% | 0.1% | 100.0% |
| 1 | Navarasa 2.0 7B india | 20.7% (17.8-23.7) | +3.6% | 46.6% | 60.0% | 0.4% | 94.0% |
| 2 | Krutrim-1 7B india | 14.4% (11.6-17.5) | +0.9% | 55.7% | 71.1% | 0.0% | 99.9% |
| 2 | Param-1 2.9B india | 14.3% (11.9-16.9) | +1.5% | 57.8% | 50.8% | 0.4% | 99.8% |
| 2 | Sarvam-1 2B india | 11.6% (9.1-14.8) | -0.1% | 60.1% | 42.2% | 0.9% | 99.8% |
| 2 | OpenHathi 7B india | 11.2% (8.7-14.2) | -0.4% | 57.8% | 58.7% | 0.1% | 96.3% |
| 2 | Airavata 7B india | 7.9% (5.6-10.4) | +1.1% | 26.7% | 14.5% | 37.4% | 100.0% |
| 2 | Param-1 7B india | 4.8% (3.2-6.6) | -0.3% | 33.2% | 19.4% | 31.6% | 99.5% |
| 2 | Gemma 3 12B INT4 global | 4.6% (2.4-7.2) | +29.6% | 20.7% | 25.8% | 46.5% | 99.7% |
Every model, every category, no truncation — ranked top to bottom by Bharat Score. Fill shade scales with accuracy so the weak column (land units, almost throughout) reads at a glance; the exact percentage is always printed alongside it.
| model | numerals | weights/vol | land units | seasons | fiscal yr | schemes | identifiers |
|---|---|---|---|---|---|---|---|
| Qwen3.6 27B global | 77.9 | 61.1 | 39.4 | 48.9 | 52.9 | 41.9 | 50.0 |
| gpt-oss-20b global | 78.7 | 57.6 | 34.9 | 41.2 | 46.7 | 25.4 | 43.3 |
| Sarvam-M 24B india | 46.4 | 55.8 | 21.5 | 54.6 | 49.0 | 36.5 | 52.2 |
| Gemma 3 27B global | 54.0 | 54.9 | 26.4 | 38.5 | 50.6 | 36.0 | 32.2 |
| Qwen3-VL 8B global | 55.4 | 39.4 | 31.2 | 42.1 | 39.8 | 29.7 | 36.7 |
| Mistral Small 3.1 24B global | 50.6 | 59.0 | 26.1 | 42.9 | 32.2 | 31.1 | 30.0 |
| Krutrim-2 12B india | 49.2 | 39.6 | 24.5 | 44.7 | 42.9 | 39.9 | 30.0 |
| Phi-4 14B global | 46.1 | 53.5 | 24.2 | 42.1 | 46.1 | 32.2 | 26.7 |
| Sarvam 30B india | 40.7 | 51.2 | 28.8 | 51.1 | 33.1 | 27.9 | 26.7 |
| Gemma 3 12B global | 48.8 | 40.7 | 20.8 | 41.9 | 37.8 | 29.0 | 26.7 |
| Llama 3.1 8B Instruct global | 32.0 | 25.7 | 16.8 | 39.6 | 38.2 | 35.1 | 40.0 |
| Qwen2.5 7B global | 44.6 | 37.7 | 17.5 | 32.4 | 34.5 | 16.0 | 33.3 |
| Gemma 3 4B global | 32.9 | 24.3 | 12.5 | 34.6 | 31.2 | 18.2 | 23.3 |
| Navarasa 2.0 7B india | 26.1 | 21.5 | 10.6 | 25.1 | 30.4 | 17.6 | 13.3 |
| Krutrim-1 7B india | 17.1 | 3.5 | 5.8 | 28.0 | 13.5 | 16.2 | 16.7 |
| Param-1 2.9B india | 15.2 | 11.1 | 4.0 | 25.8 | 15.3 | 15.5 | 13.3 |
| Sarvam-1 2B india | 16.2 | 2.8 | 2.4 | 18.5 | 12.8 | 14.9 | 13.3 |
| OpenHathi 7B india | 13.1 | 4.9 | 2.2 | 15.4 | 16.7 | 8.8 | 17.8 |
| Airavata 7B india | 6.9 | 4.9 | 1.4 | 12.6 | 11.8 | 4.0 | 13.3 |
| Param-1 7B india | 7.4 | 0.7 | 0.0 | 6.0 | 4.1 | 5.4 | 10.0 |
| Gemma 3 12B INT4 global | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 12.2 | 20.0 |
Paired bootstrap test between each rank and the one directly below it, Holm-Bonferroni corrected across all 20 tests at once (0.05 family-wise alpha) — the "significant" column, not the raw p-value alone, is the real "these two models differ" claim. 7 of 20 adjacent pairs survive correction.
| rank | next rank | mean diff | p-value | Holm-corrected |
|---|---|---|---|---|
| Qwen3.6 27B | gpt-oss-20b | +5.8pp | <0.001 | significant |
| gpt-oss-20b | Sarvam-M 24B | +5.0pp | 0.001 | significant |
| Sarvam-M 24B | Gemma 3 27B | +0.7pp | 0.669 | not significant |
| Gemma 3 27B | Qwen3-VL 8B | +2.6pp | 0.068 | not significant |
| Qwen3-VL 8B | Mistral Small 3.1 24B | +0.4pp | 0.781 | not significant |
| Mistral Small 3.1 24B | Krutrim-2 12B | -0.0pp | 0.993 | not significant |
| Krutrim-2 12B | Phi-4 14B | -0.1pp | 0.957 | not significant |
| Phi-4 14B | Sarvam 30B | +1.6pp | 0.334 | not significant |
| Sarvam 30B | Gemma 3 12B | +1.9pp | 0.26 | not significant |
| Gemma 3 12B | Llama 3.1 8B Instruct | +5.2pp | 0.001 | significant |
| Llama 3.1 8B Instruct | Qwen2.5 7B | +0.3pp | 0.871 | not significant |
| Qwen2.5 7B | Gemma 3 4B | +5.2pp | <0.001 | significant |
| Gemma 3 4B | Navarasa 2.0 7B | +4.1pp | 0.001 | significant |
| Navarasa 2.0 7B | Krutrim-1 7B | +7.3pp | <0.001 | significant |
| Krutrim-1 7B | Param-1 2.9B | -0.0pp | 0.966 | not significant |
| Param-1 2.9B | Sarvam-1 2B | +2.9pp | 0.009 | not significant |
| Sarvam-1 2B | OpenHathi 7B | +1.0pp | 0.375 | not significant |
| OpenHathi 7B | Airavata 7B | +3.4pp | 0.001 | significant |
| Airavata 7B | Param-1 7B | +2.9pp | 0.004 | not significant |
| Param-1 7B | Gemma 3 12B INT4 | +2.0pp | 0.011 | not significant |
Every item is asked twice per model: once demanding a strict JSON-shaped answer (R1) and once allowing a natural-language response (R2). Roster average: R1 30.1%, R2 26.3% (-3.7pp).
| model | R1 (strict JSON) | R2 (natural language) | R2 − R1 |
|---|---|---|---|
| Qwen3.6 27B | 63.5% | 42.8% | -20.7pp |
| gpt-oss-20b | 53.8% | 39.9% | -14.0pp |
| Sarvam-M 24B | 53.6% | 36.6% | -17.0pp |
| Gemma 3 27B | 47.7% | 35.9% | -11.8pp |
| Qwen3-VL 8B | 39.8% | 38.6% | -1.2pp |
| Mistral Small 3.1 24B | 39.0% | 38.6% | -0.4pp |
| Krutrim-2 12B | 40.0% | 37.4% | -2.7pp |
| Phi-4 14B | 44.2% | 33.2% | -11.0pp |
| Sarvam 30B | 40.7% | 33.5% | -7.2pp |
| Gemma 3 12B | 37.6% | 32.6% | -5.0pp |
| Llama 3.1 8B Instruct | 35.1% | 29.9% | -5.2pp |
| Qwen2.5 7B | 29.7% | 32.0% | +2.3pp |
| Gemma 3 4B | 25.5% | 25.1% | -0.4pp |
| Navarasa 2.0 7B | 22.5% | 18.9% | -3.6pp |
| Krutrim-1 7B | 11.3% | 17.5% | +6.2pp |
| Param-1 2.9B | 14.0% | 14.7% | +0.7pp |
| Sarvam-1 2B | 10.3% | 12.8% | +2.4pp |
| OpenHathi 7B | 11.5% | 11.1% | -0.4pp |
| Airavata 7B | 4.0% | 11.8% | +7.8pp |
| Param-1 7B | 1.7% | 7.9% | +6.2pp |
| Gemma 3 12B INT4 | 6.5% | 2.7% | -3.8pp |
A pre-registered secondary cut: Sarvam, Krutrim, Airavata/Navarasa/OpenHathi (AI4Bharat), and Param (BharatGenAI/CDAC) bucketed as India-built/tuned (10 models) against every other model in the roster as a global open-weight baseline (11 models). This is descriptive, not a controlled comparison — family, size, and training recipe all vary within each bucket — and with roughly 10–11 models per side, expect real sampling noise in a per-bucket mean.
| category | india mean | global mean | gap (global−india) |
|---|---|---|---|
| Indian numeral system | 23.8% | 47.4% | +23.5pp |
| Weights, volumes & informal measures | 19.6% | 41.3% | +21.7pp |
| Land units by state | 10.1% | 22.7% | +12.6pp |
| Agricultural seasons & crop calendars | 28.2% | 36.8% | +8.6pp |
| Fiscal year & date conventions | 23.0% | 37.3% | +14.3pp |
| Government scheme literacy | 18.7% | 27.9% | +9.2pp |
| Structural/format identifiers | 20.7% | 32.9% | +12.3pp |
Lowest mean accuracy across every model that scored the item. All 10 of these items were answered correctly by zero of the 21 models scored.
| category | prompt | gold | mean accuracy | n models |
|---|---|---|---|---|
| Indian numeral system | What is "paune crore" expressed as a plain number of rupees? | 7500000.0 | 0.0% | 21 |
| Weights, volumes & informal measures | A jeweller's inventory lists 2 chhatak of gold. Express this in grams, rounded to two decimals. | 116.64 | 0.0% | 21 |
| Weights, volumes & informal measures | A goldsmith's inventory is recorded as 3 tola 2 chhatak. Express the total in grams, rounded to two decimals. | 151.63 | 0.0% | 21 |
| Weights, volumes & informal measures | One tola equals how many chhatak? | 0.2 | 0.0% | 21 |
| Weights, volumes & informal measures | A warehouse holds 250 cotton bales. Express the total weight in tonnes, rounded to two decimals. | 42.5 | 0.0% | 21 |
| Weights, volumes & informal measures | Which weighs more: 1 cotton bale or 1 jute bale, and by how many kilograms? | 10.0 | 0.0% | 21 |
| Land units by state | A plot in Uttar Pradesh is recorded as 5 pucca bigha. Express the area in acres, rounded to two decimals. | 3.125 | 0.0% | 21 |
| Land units by state | A landholding in Varanasi district, Uttar Pradesh is recorded as 3 pucca bigha 8 biswa. Express the area in acres, rounded to two decimals. | 2.13 | 0.0% | 21 |
| Land units by state | In and around Patna, Bihar, one kattha is approximately how many square feet? | 1361.25 | 0.0% | 21 |
| Land units by state | In Muzaffarpur, Bihar, one kattha is stated as how many square feet? | 1500 | 0.0% | 21 |
Order-of-magnitude slips — a lakh/crore-scale value parsed off by a large, plausible-looking ratio (not degenerate repetition):
Overconfidence — an ambiguous, underspecified-premise item answered with a single confident value instead of a hedge:
Unit drop — the numeric value is right, the explicit unit is missing or wrong:
Every response is graded deterministically — code, not a model, decides right/wrong. The grader is picked per item off its grader field, one of 8 fixed graders keyed to the item's answer type; there is no LLM in the grading loop at all.
| answer type | grader | how it decides |
|---|---|---|
| numeric | numeric_v1 | parses a number out of the response (JSON field on R1, best-effort extraction on R2) and checks it against gold within the item's own tolerance — exact, or relative/absolute within a stated margin |
| numeric_with_unit | numeric_unit_v1 | same numeric tolerance check, plus the response's unit string must match the gold unit (tracked separately as Unit Discipline) |
| date | date_v1 | parses a calendar date out of the response and compares it to the gold date exactly |
| date_range | date_range_v1 | parses a start/end date pair and compares both ends to gold |
| month_set | month_set_v1 | parses the set of months named in the response and compares it to gold's month set |
| enum | enum_v1 | matches the response against gold plus its accepted_aliases list (case/punctuation-normalized) |
| string_normalized | string_normalized_v1 | normalizes both response and gold the same way (lowercase/strip, or a stricter digit-grouping/word-form mode per item) before comparing |
| clarification | clarification_v1 | checks for a hedge (e.g. "depends on state," "could be X or Y") rather than one confident value; a single unhedged value on a deliberately ambiguous item is marked incorrect and flagged overconfident |