WE BENCHMARKED 21 MODELS · BKP-500 FULL RUNS (R1+R2, n=3 SAMPLES) · 1084 ITEMS EACH
~/research/bkp500-2026
BKP-500, the 2026 model run report

We ran 21 models against BKP-500 — 552 India-specific core items across 7 categories (numerals, weights/volumes, land units, agricultural seasons, fiscal-year conventions, government schemes, structural identifiers), each paired 1:1 against an international control twin (532 control items), under two prompt regimes and 3 repeated samples per item. Qwen3.6 27B leads the roster at 53.2% Bharat Score; Sarvam-M 24B is the strongest India-built/tuned model at 45.1%. The single weakest spot for the roster as a whole is Land units by state — the worst-scoring category for 15 of 21 models.

~/research/bkp500-2026

BKP-500, the 2026 model run.

We scored 21 models · 1084 items (552 core + 532 control) · 2 prompt regimes × 3 samples = 6504 rows/model · run on the BKP evaluation pipeline

summary

BKP-500 tests whether a model actually knows India-specific conventions — lakh/crore numerals, informal weights (tola, seer, chhatak), land units that vary by state, kharif/rabi/zaid crop calendars, the April–March fiscal year, and government-scheme facts — rather than defaulting to the international convention that dominates most pretraining data. Every core item is paired with a matched control item that asks the equivalent question under an international/self-contained convention, so a model's Locale Gap Delta (control accuracy minus India accuracy) isolates a genuine India-specific knowledge gap from general capability.

Qwen3.6 27B leads the roster at 53.2% Bharat Score (95% CI 48.9–57.3%), a global open-weight model, not an India-built one. Sarvam-M 24B is the strongest India-built/tuned model at 45.1%, landing at rank 3 overall. Across the full roster, global open-weight models average 35.2% against India-built models' 20.6% — a 14.6-point gap, though with only 10 and 11 models per bucket that comparison's own confidence interval is wide (see the bucket comparison below). Land units by state is the hardest category for the roster as a whole: the single lowest-scoring category for 15 of 21 models, well ahead of any other category.

key insights

  1. The top three ranks are all real, not noise — then an 8-model plateau.

    Qwen3.6 27B's lead over gpt-oss-20b (6.3 points of Bharat Score) and gpt-oss-20b's lead over Sarvam-M 24B (1.7 points) both survive Holm-Bonferroni correction across every adjacent-rank test at once (0.05 family-wise alpha). From rank 3 (Sarvam-M 24B) through rank 10 (Gemma 3 12B), no adjacent pair is statistically distinguishable — a genuine 8-model plateau. It doesn't hold for the rest of tier 1, though: 3 more Holm-significant breaks punctuate the remaining ranks (Gemma 3 12B→Llama 3.1 8B Instruct, Qwen2.5 7B→Gemma 3 4B, Gemma 3 4B→Navarasa 2.0 7B) before the tier 1/tier 2 cliff itself (see below).

  2. An India-built model holds rank 3 outright.

    Sarvam-M 24B scores 45.1%, tier 1 — 1.7 points behind #2 gpt-oss-20b (a Holm-significant gap), 3.3 points ahead of #4 Gemma 3 27B (not a Holm-significant gap). It shares tier 1's upper half with 1 other India-built/tuned model (Krutrim-2 12B).

  3. Land units by state breaks almost every model, regardless of scale or origin.

    Land units by state is the worst-scoring category for 15 of 21 models — the next most common weak spot is a distant runner-up. State-by-state variation in land-unit definitions (a bigha in Punjab isn't a bigha in Bihar) appears to be genuinely harder for every model family we tested than any other category in the corpus.

  4. The tier 1/tier 2 cliff is a real, significant gap.

    Navarasa 2.0 7B (20.7%, tier 1) and Krutrim-1 7B (14.4%, tier 2) sit only 6.3 points apart in raw score, but that gap is Holm-significant (p<0.001) — tier 2 isn't an artifact of the tiering method, it's a real capability floor 7 of 21 models fall below.

  5. Locale Gap Delta doesn't move in lockstep with overall accuracy.

    Gemma 3 12B INT4 shows the largest control-favoring gap in the roster by far (+29.6 points, next-highest is under 4) — but it's also the roster's lowest-scoring model overall (4.6%, rank 21 of 21), so this extreme reading is more likely a symptom of that model's broader breakdown than a clean locale-specific signal. Qwen3.6 27B is a cleaner example of the gap moving independently of accuracy: it's rank 1 of 21 by Bharat Score (53.2%), yet it's actually 14.0 points easier on India-framed items than on the matched international control — the reverse of the expected direction, from a model that's clearly not broken.

  6. Wrong answers are consistent, not noisy.

    Consistency — agreement across a model's own repeated samples — sits at 87.5–100.0% across the entire roster, regardless of Bharat Score. A model scoring 5% and one scoring 50% are both answering the same way every time they're asked — low accuracy here means confidently, repeatably wrong, not random guessing.

Also worth noting

  • Samples are near-deterministic in practice: 78.9% of multi-sample groups (35,915 of 45,528) returned byte-identical text across all 3 repeated samples at temperature 0, as configured.
  • 1,086 of 1,086 items in the current corpus (100.0%) — not the same 1084 items the roster above was actually scored on, since 2 control item(s) were added after that run — have been through a human review pass: a reviewer flagged 123 with a gold-value problem, 104 of those corrections are already applied to the live corpus, and 19 remain open (mostly `clarification_required` items where the reviewer proposed a single scalar value against a deliberately multi-source-conflicting convention — left as clarification pending resolution, not auto-overridden). This single-reviewer QA pass is a different, lighter bar than the project's stricter >=2-independent-annotator adjudication gate for dev/test eligibility, which hasn't run yet (see methodology below) — every item is still split="draft" in that stricter sense.

findings so far

01

INT4-quantizing Gemma 3 12B doesn't cost a few points — it erases the model.

Gemma 3 12B scores 35.1% (rank 10 of 21). The same model, quantized to INT4, scores 4.6% — dead last, rank 21 of 21, below even Gemma 3 4B (25.3%). Its refusal rate jumps to 46.5% and it scores a flat 0.0% on 5 of the 7 categories — this reads as a broken quantization, not a graceful capability trade-off.

02

The 3 best models take the biggest hit from strict JSON formatting.

Qwen3.6 27B loses 20.7 points going from natural-language (R2) to strict-JSON (R1) answers, Sarvam-M 24B loses 17.0, gpt-oss-20b loses 14.0 — and these are exactly the roster's top 3 ranks by Bharat Score. Meanwhile Airavata 7B actually gains 7.8 points under strict JSON. Format compliance and raw capability pull in opposite directions here: the strongest models seem to lose more to answer-extraction friction, not less.

03

Best-calibrated on ambiguity isn't the best-scoring model overall.

Sarvam 30B leads the roster on Ambiguity Handling (69.4%) with a low Overconfidence rate (8.3%), despite ranking 9 of 21 overall. At the other extreme, Param-1 7B and Gemma 3 12B INT4 both score 0% Ambiguity Handling — never hedging on a deliberately underspecified item — but for very different reasons: Param-1 7B answers confidently and wrongly (58.3% overconfidence), while Gemma 3 12B INT4's own 0.0% overconfidence rate suggests it's mostly just not answering at all.

04

The best unit discipline in the roster belongs to India-built models.

Sarvam 30B, Sarvam-M 24B and Krutrim-1 7B post the roster's top 3 Unit Discipline scores (75.4%, 74.7%, 71.1%) — all three India-built/tuned, despite mixed overall ranks (9, 3, 15 of 21). Airavata 7B and Param-1 7B trail furthest behind (14.5%, 19.4%) — both are also among the roster's highest-refusal models, so a chunk of their missing units is likely missing answers, not wrong ones.

05

Qwen's generation-then-scale jump is the cleanest improvement story in the roster.

Qwen2.5 7B (30.9%, rank 12) → Qwen3-VL 8B (39.2%, rank 5, +8.3pp at near-identical size) is a generational jump, not a scale one. Qwen3.6 27B (53.2%, rank 1) then adds real scale on top, landing as the roster's outright leader. Every step in this family's chain moved in the same direction — no other model family in the roster shows as clean a progression.

06

Indian numeral system shows the widest India-built-vs-global gap; Agricultural seasons & crop calendars the narrowest.

Global open-weight models lead India-built/tuned models by 23.5 points on Indian numeral system and 21.7 on Weights, volumes & informal measures — the two categories built on precise lakh/crore-scale arithmetic. The gap narrows to 8.6 points on Agricultural seasons & crop calendars, the smallest of any category. The India-built/tuned bucket doesn't trail evenly across categories: it closes ground fastest on categories that reward cultural/domain familiarity over precise numeric conversion, and falls furthest behind exactly where scale arithmetic dominates.

league table

Ranked by Bharat Score (macro-mean accuracy across all 7 categories), tiered by chained bootstrap-CI overlap. Locale Gap Delta = control-item accuracy minus India-item accuracy on matched pairs; positive means the international framing was easier for that model, negative means the India-specific framing was.

tiermodelbharat score (95% CI)locale gap Δoom error rateunit disciplinerefusal rateconsistency
1Qwen3.6 27B global 53.2% (48.9-57.3) -14.0% 23.0% 69.8% 0.7% 100.0%
1gpt-oss-20b global 46.8% (43.0-50.9) -9.4% 25.8% 65.7% 0.1% 89.3%
1Sarvam-M 24B india 45.1% (40.9-49.1) -8.6% 32.8% 74.7% 3.6% 92.6%
1Gemma 3 27B global 41.8% (38.3-45.8) -5.8% 27.8% 68.6% 0.1% 99.6%
1Qwen3-VL 8B global 39.2% (35.1-43.4) -1.8% 28.7% 66.9% 0.0% 97.6%
1Mistral Small 3.1 24B global 38.8% (35.0-42.9) -3.5% 26.7% 61.0% 0.4% 98.2%
1Krutrim-2 12B india 38.7% (34.6-43.3) -5.8% 32.5% 66.4% 0.2% 99.6%
1Phi-4 14B global 38.7% (34.8-43.0) -4.0% 29.2% 68.1% 3.7% 99.0%
1Sarvam 30B india 37.1% (33.5-40.9) -6.3% 40.0% 75.4% 2.3% 87.5%
1Gemma 3 12B global 35.1% (31.3-39.4) -2.8% 32.4% 66.7% 0.0% 99.7%
1Llama 3.1 8B Instruct global 32.5% (28.5-36.6) -3.1% 40.4% 63.5% 2.2% 100.0%
1Qwen2.5 7B global 30.9% (27.1-34.8) +1.2% 31.6% 67.2% 0.1% 98.1%
1Gemma 3 4B global 25.3% (22.1-28.8) -2.0% 41.3% 60.9% 0.1% 100.0%
1Navarasa 2.0 7B india 20.7% (17.8-23.7) +3.6% 46.6% 60.0% 0.4% 94.0%
2Krutrim-1 7B india 14.4% (11.6-17.5) +0.9% 55.7% 71.1% 0.0% 99.9%
2Param-1 2.9B india 14.3% (11.9-16.9) +1.5% 57.8% 50.8% 0.4% 99.8%
2Sarvam-1 2B india 11.6% (9.1-14.8) -0.1% 60.1% 42.2% 0.9% 99.8%
2OpenHathi 7B india 11.2% (8.7-14.2) -0.4% 57.8% 58.7% 0.1% 96.3%
2Airavata 7B india 7.9% (5.6-10.4) +1.1% 26.7% 14.5% 37.4% 100.0%
2Param-1 7B india 4.8% (3.2-6.6) -0.3% 33.2% 19.4% 31.6% 99.5%
2Gemma 3 12B INT4 global 4.6% (2.4-7.2) +29.6% 20.7% 25.8% 46.5% 99.7%
OOM error rate = share of numeric answers off by ≥10× (a lakh/crore-scale slip). Unit discipline = share of numeric_with_unit answers that included an explicit, correct unit. n_resamples=2,000 for every bootstrap CI on this page (the published leaderboard uses 10,000 for the publication-grade run).

per-category accuracy — all 21 models × 7 categories

Every model, every category, no truncation — ranked top to bottom by Bharat Score. Fill shade scales with accuracy so the weak column (land units, almost throughout) reads at a glance; the exact percentage is always printed alongside it.

modelnumeralsweights/volland unitsseasonsfiscal yrschemesidentifiers
Qwen3.6 27B global77.961.139.448.952.941.950.0
gpt-oss-20b global78.757.634.941.246.725.443.3
Sarvam-M 24B india46.455.821.554.649.036.552.2
Gemma 3 27B global54.054.926.438.550.636.032.2
Qwen3-VL 8B global55.439.431.242.139.829.736.7
Mistral Small 3.1 24B global50.659.026.142.932.231.130.0
Krutrim-2 12B india49.239.624.544.742.939.930.0
Phi-4 14B global46.153.524.242.146.132.226.7
Sarvam 30B india40.751.228.851.133.127.926.7
Gemma 3 12B global48.840.720.841.937.829.026.7
Llama 3.1 8B Instruct global32.025.716.839.638.235.140.0
Qwen2.5 7B global44.637.717.532.434.516.033.3
Gemma 3 4B global32.924.312.534.631.218.223.3
Navarasa 2.0 7B india26.121.510.625.130.417.613.3
Krutrim-1 7B india17.13.55.828.013.516.216.7
Param-1 2.9B india15.211.14.025.815.315.513.3
Sarvam-1 2B india16.22.82.418.512.814.913.3
OpenHathi 7B india13.14.92.215.416.78.817.8
Airavata 7B india6.94.91.412.611.84.013.3
Param-1 7B india7.40.70.06.04.15.410.0
Gemma 3 12B INT4 global0.00.00.00.00.012.220.0
Land units by state is the single weakest category for 15 of 21 models in the full roster.

pairwise significance — every adjacent leaderboard rank

Paired bootstrap test between each rank and the one directly below it, Holm-Bonferroni corrected across all 20 tests at once (0.05 family-wise alpha) — the "significant" column, not the raw p-value alone, is the real "these two models differ" claim. 7 of 20 adjacent pairs survive correction.

ranknext rankmean diffp-valueHolm-corrected
Qwen3.6 27Bgpt-oss-20b+5.8pp<0.001significant
gpt-oss-20bSarvam-M 24B+5.0pp0.001significant
Sarvam-M 24BGemma 3 27B+0.7pp0.669not significant
Gemma 3 27BQwen3-VL 8B+2.6pp0.068not significant
Qwen3-VL 8BMistral Small 3.1 24B+0.4pp0.781not significant
Mistral Small 3.1 24BKrutrim-2 12B-0.0pp0.993not significant
Krutrim-2 12BPhi-4 14B-0.1pp0.957not significant
Phi-4 14BSarvam 30B+1.6pp0.334not significant
Sarvam 30BGemma 3 12B+1.9pp0.26not significant
Gemma 3 12BLlama 3.1 8B Instruct+5.2pp0.001significant
Llama 3.1 8B InstructQwen2.5 7B+0.3pp0.871not significant
Qwen2.5 7BGemma 3 4B+5.2pp<0.001significant
Gemma 3 4BNavarasa 2.0 7B+4.1pp0.001significant
Navarasa 2.0 7BKrutrim-1 7B+7.3pp<0.001significant
Krutrim-1 7BParam-1 2.9B-0.0pp0.966not significant
Param-1 2.9BSarvam-1 2B+2.9pp0.009not significant
Sarvam-1 2BOpenHathi 7B+1.0pp0.375not significant
OpenHathi 7BAiravata 7B+3.4pp0.001significant
Airavata 7BParam-1 7B+2.9pp0.004not significant
Param-1 7BGemma 3 12B INT4+2.0pp0.011not significant
Mean diff = mean per-item core-accuracy difference (higher-ranked minus next-ranked model), bootstrapped over every core item common to both — NOT the same number as subtracting the two models' Bharat Score column above. Bharat Score is a macro-mean across the 7 categories (each category weighted equally); this diff pools every item's own correct/incorrect flag directly, unweighted by category, so a large, low-scoring category like land units pulls it more than it pulls the equal-weighted Bharat Score. Both are legitimate statistics answering different questions — this one is what the significance test itself resamples over, so its p-value/Holm-corrected verdict is the number to trust for "do these two differ," while the league table's Bharat Score is the one to trust for "by how much." A wide run of "not significant" rows (as in the middle of tier 1 here) means those ranks are a real statistical plateau, not a meaningfully fine-grained order.

R1 (strict JSON) vs. R2 (natural language) — Bharat Score by regime

Every item is asked twice per model: once demanding a strict JSON-shaped answer (R1) and once allowing a natural-language response (R2). Roster average: R1 30.1%, R2 26.3% (-3.7pp).

modelR1 (strict JSON)R2 (natural language)R2 − R1
Qwen3.6 27B63.5%42.8%-20.7pp
gpt-oss-20b53.8%39.9%-14.0pp
Sarvam-M 24B53.6%36.6%-17.0pp
Gemma 3 27B47.7%35.9%-11.8pp
Qwen3-VL 8B39.8%38.6%-1.2pp
Mistral Small 3.1 24B39.0%38.6%-0.4pp
Krutrim-2 12B40.0%37.4%-2.7pp
Phi-4 14B44.2%33.2%-11.0pp
Sarvam 30B40.7%33.5%-7.2pp
Gemma 3 12B37.6%32.6%-5.0pp
Llama 3.1 8B Instruct35.1%29.9%-5.2pp
Qwen2.5 7B29.7%32.0%+2.3pp
Gemma 3 4B25.5%25.1%-0.4pp
Navarasa 2.0 7B22.5%18.9%-3.6pp
Krutrim-1 7B11.3%17.5%+6.2pp
Param-1 2.9B14.0%14.7%+0.7pp
Sarvam-1 2B10.3%12.8%+2.4pp
OpenHathi 7B11.5%11.1%-0.4pp
Airavata 7B4.0%11.8%+7.8pp
Param-1 7B1.7%7.9%+6.2pp
Gemma 3 12B INT46.5%2.7%-3.8pp
A model that scores much lower under R1 is likely losing points to answer-extraction friction (malformed JSON, missing fields) rather than to actually not knowing the fact — a large gap in either direction is worth checking against the raw responses before reading it as a knowledge difference.

india-built vs. global open-weight

A pre-registered secondary cut: Sarvam, Krutrim, Airavata/Navarasa/OpenHathi (AI4Bharat), and Param (BharatGenAI/CDAC) bucketed as India-built/tuned (10 models) against every other model in the roster as a global open-weight baseline (11 models). This is descriptive, not a controlled comparison — family, size, and training recipe all vary within each bucket — and with roughly 10–11 models per side, expect real sampling noise in a per-bucket mean.

India-built (10)
20.6
Global open-weight (11)
35.2
categoryindia meanglobal meangap (global−india)
Indian numeral system23.8%47.4%+23.5pp
Weights, volumes & informal measures19.6%41.3%+21.7pp
Land units by state10.1%22.7%+12.6pp
Agricultural seasons & crop calendars28.2%36.8%+8.6pp
Fiscal year & date conventions23.0%37.3%+14.3pp
Government scheme literacy18.7%27.9%+9.2pp
Structural/format identifiers20.7%32.9%+12.3pp
Bucket-level Bharat Score 95% CI (bootstrapped over models within the bucket, not items): india 12.4–29.4%, global 27.6–41.6%. The intervals are wide at this sample size (10/11 models) — that width is itself the honest finding, not a gap to paper over.

hardest items

Lowest mean accuracy across every model that scored the item. All 10 of these items were answered correctly by zero of the 21 models scored.

categorypromptgoldmean accuracyn models
Indian numeral systemWhat is "paune crore" expressed as a plain number of rupees?7500000.00.0%21
Weights, volumes & informal measuresA jeweller's inventory lists 2 chhatak of gold. Express this in grams, rounded to two decimals.116.640.0%21
Weights, volumes & informal measuresA goldsmith's inventory is recorded as 3 tola 2 chhatak. Express the total in grams, rounded to two decimals.151.630.0%21
Weights, volumes & informal measuresOne tola equals how many chhatak?0.20.0%21
Weights, volumes & informal measuresA warehouse holds 250 cotton bales. Express the total weight in tonnes, rounded to two decimals.42.50.0%21
Weights, volumes & informal measuresWhich weighs more: 1 cotton bale or 1 jute bale, and by how many kilograms?10.00.0%21
Land units by stateA plot in Uttar Pradesh is recorded as 5 pucca bigha. Express the area in acres, rounded to two decimals.3.1250.0%21
Land units by stateA landholding in Varanasi district, Uttar Pradesh is recorded as 3 pucca bigha 8 biswa. Express the area in acres, rounded to two decimals.2.130.0%21
Land units by stateIn and around Patna, Bihar, one kattha is approximately how many square feet?1361.250.0%21
Land units by stateIn Muzaffarpur, Bihar, one kattha is stated as how many square feet?15000.0%21

failure gallery

Order-of-magnitude slips — a lakh/crore-scale value parsed off by a large, plausible-looking ratio (not degenerate repetition):

$ Param-1 2.9B, Indian numeral system — "₹1.2 lakh crore is equivalent to how many USD trillion, given lakh crore == trillion in scale (ignore exchange rate)?" — gold: 1.2 trillion. model said: "{"value": 1200000000000, "unit": "trillion"}" (off by ~1,000,000,000,000×)
$ Param-1 7B, Indian numeral system — "Capital expenditure is ₹10 lakh crore, which is 3.3% of GDP. Express GDP in USD trillion at ₹83/USD (treat 1 lakh crore == 1 trillion in scale), rounded to two decimals." — gold: 3.65 trillion. model said: "So, the answer is: \[ \boxed{10,000,000,000,000}" (off by ~1,000,000,000,000×)
$ Param-1 7B, Indian numeral system — "Revenue of ₹4.5 crore grew by 18%. Express the new revenue in lakh." — gold: 531.0 lakh. model said: "Solution with explanation: {"value": 450000000000000000000" (off by ~1,000,000,000,000×)
$ Param-1 7B, Indian numeral system — "A ₹3.6 crore fund grew by one-third. Express the new value in lakh." — gold: 480.0 lakh. model said: "Answer: 3600000000000000000000000000" (off by ~1,000,000,000,000×)

Overconfidence — an ambiguous, underspecified-premise item answered with a single confident value instead of a hedge:

$ Airavata 7B, Weights, volumes & informal measures — "Convert 10 maund to kilograms." — ambiguous/underspecified premise, no hedge. model said: "Convert 10 maund to kilograms."
$ Airavata 7B, Weights, volumes & informal measures — "Express 1 candy in kilograms." — ambiguous/underspecified premise, no hedge. model said: "1 kilogram of candy is equal to 1,000 grams."
$ Airavata 7B, Weights, volumes & informal measures — "Express 2 candy in kilograms." — ambiguous/underspecified premise, no hedge. model said: "The candy Express 2 is a popular brand of candy in the United States. It is a type of candy that is made from sugar, corn syrup, and other ingredients. The exact ingredients and proportions of the candy can vary dependin"

Unit drop — the numeric value is right, the explicit unit is missing or wrong:

$ Airavata 7B, Indian numeral system — "Express 650 crore in millions." — gold: 6500.0 million (value right, unit dropped). model said: "650 करोड़ रुपये का अर्थ है 6500 करोड़ रुपये।"
$ Airavata 7B, Weights, volumes & informal measures — "Production of oilseeds is reported at 2.1 million tonnes. Express this in lakh tonnes." — gold: 21.0 lakh tonnes (value right, unit dropped). model said: "21 लाख"
$ Gemma 3 12B, Weights, volumes & informal measures — "The MSP for tur is Rs.6620 per quintal. Express this in Rs. per kg." — gold: 66.2 rupees per kg (value right, unit dropped). model said: "{"value": 66.2, "unit": "Rs."}"

methodology

$ protocol    →  3 samples × regimes R1, R2 at temperature 0.0 per item; 552 core items paired 1:1 against 534 matched control items in the live corpus (18 core items are unpaired by design — clarification-type items with no international twin, scored via Ambiguity Handling/Overconfidence instead). 2 of those 534 matched controls were added to the corpus after this run completed and were never sent to any model — excluded from the 532-control count above, which is what the roster was actually scored on.
$ statistics  →  every CI on this page (Bharat Score, per-category, Locale Gap Delta) is bootstrapped over one value per item, not per raw response row, to avoid treating 3 repeated samples of one item as 3 independent items. Adjacent-rank significance tests are Holm-Bonferroni corrected across the whole family at 0.05 alpha — the `holm_significant` field, not each test's raw p-value, is the real "these two models differ" claim
$ human review →  100.0% coverage — 1,086 of 1,086 core+control items in the current corpus (2 more than the 1084 the roster was actually scored on) have been through a single-reviewer QA pass (123 flagged, 104 corrections applied, 19 left open as genuine unresolved disagreements). This is not the same as the project's stricter ≥2-independent-human-annotator adjudication gate for dev/test eligibility, which has not run — 0 of 1,086 items are adjudicated under that gate, and every item is still split="draft" in that sense
$ pinned     →  4 of 22 models in the run registry carry an explicit HF revision pin; the rest resolve to whatever the backend served on the run date. (1 registry model — llama4-scout — produced zero graded rows (a crashed run, per its own run log) and is excluded entirely, not just partially, from the 21 models scored above.)

how answers are graded

Every response is graded deterministically — code, not a model, decides right/wrong. The grader is picked per item off its grader field, one of 8 fixed graders keyed to the item's answer type; there is no LLM in the grading loop at all.

answer typegraderhow it decides
numericnumeric_v1parses a number out of the response (JSON field on R1, best-effort extraction on R2) and checks it against gold within the item's own tolerance — exact, or relative/absolute within a stated margin
numeric_with_unitnumeric_unit_v1same numeric tolerance check, plus the response's unit string must match the gold unit (tracked separately as Unit Discipline)
datedate_v1parses a calendar date out of the response and compares it to the gold date exactly
date_rangedate_range_v1parses a start/end date pair and compares both ends to gold
month_setmonth_set_v1parses the set of months named in the response and compares it to gold's month set
enumenum_v1matches the response against gold plus its accepted_aliases list (case/punctuation-normalized)
string_normalizedstring_normalized_v1normalizes both response and gold the same way (lowercase/strip, or a stricter digit-grouping/word-form mode per item) before comparing
clarificationclarification_v1checks for a hedge (e.g. "depends on state," "could be X or Y") rather than one confident value; a single unhedged value on a deliberately ambiguous item is marked incorrect and flagged overconfident
If a grader can't extract anything to compare at all (garbled or empty output), it declines rather than guessing — that response lands in undetermined_rate, not in the correct or incorrect count. Every response tries the R1 (strict JSON) extraction path first regardless of which regime it was actually asked under, so a model that replies with JSON under R2 isn't penalized for it.

credit & links

BKP-500 is this project's own corpus and evaluation pipeline — not an adopted third-party benchmark. Every item, grader, and statistic on this page comes from the same evaluation pipeline that powers the canonical leaderboard, so nothing here is a separately-maintained number that can drift from what actually ran. This is the single generated report for the project — an earlier, separate statistical writeup has been folded into the sections above rather than kept as a second file.

Every item has been through a 100%-coverage human review pass (see methodology); 19 items remain flagged as open disagreements rather than mechanically resolved, and the stricter two-independent-annotator adjudication gate for dev/test eligibility hasn't run yet. Treat every figure on this page as reviewed but not yet fully adjudicated.