WE BENCHMARKED 17 MODELS · FULL 11-LANGUAGE MILU RUNS · 79,608 QUESTIONS EACH
~/benchmark/milu-2026

MILU, re-run on the 2026 models.

We scored 17 models, full 11-language runs · 79,608 questions per model · benchmark by AI4Bharat & IBM · we ran this on the Sthānika evaluation pipeline · Aug 2026

summary

We ran the full 11-language MILU test set — 79,608 questions each — on 17 models available through mid-2026 (16 league-table rows: 15 open-weight models, with Qwen3.6-27B appearing twice, thinking on and off — the other 2 models were API-only and are covered separately below). We matched the scoring protocol to each model: loglikelihood 5-shot where the harness fits, generative 0-shot where it doesn't, and paired every result with a real or reference-derived cost per 1,000 answered questions.

Qwen3.6-27B with thinking on leads the open-weight league table at 83.84% overall, ahead of Sarvam-M 24B's 82.44%. Its own thinking-off configuration scores 16.84 points lower (67.00%), the widest swing from a single setting change in the roster. Two models in our roster — Qwen3.8-Max (89.67%, ₹187.45/1K) and DeepSeek V4-Flash (79.51%, ₹2.45/1K) — were accessed only via API with no local checkpoint, so we compare them separately against closed frontier models rather than ranking them alongside models we can actually re-host ourselves. Beyond the rankings, the data show that scoring protocol is the main confounder for reasoning models, matched-size generations still gain meaningfully (Gemma 3→4, Qwen2.5→Qwen3-VL), and the low-resource gap — especially Odia — persists across almost every model regardless of scale.

key insights

  1. Alibaba's Qwen3.6-27B leads India's own best open model on our leaderboard.

    Qwen3.6-27B, run with thinking enabled, tops our open-weight leaderboard at 83.84% overall, ahead of Sarvam-M 24B's 82.44% — India's strongest open model. The catch: it isn't cheap. At ₹161.35 per 1,000 answers, it costs nearly 4× what Sarvam-M does (₹41.85/1K) — a real trade-off, not a free win.

  2. Two models tied within a point overall — but who's "ahead" flips by subject.

    OpenAI's open-weight gpt-oss-20b trails Sarvam-30B by less than a point overall (72.22% vs 72.99%). Break it down by domain and the lead flips: gpt-oss-20b actually wins Science (+2.6 points) and Social Sciences (+1.4 points). Sarvam-30B pulls furthest ahead in Law & Governance, by 5.9 points — of every foreign-vs-Indic open-model pair we checked, this is the only one where the domain-level lead changes hands at all.

  3. Alibaba's model beats India's own flagship in 9 of 10 Indian languages — and it isn't close in Odia.

    Qwen3.6-27B doesn't just edge past Sarvam-M 24B on the overall score — it out-scores Sarvam-M in 9 of the 10 Indic languages we tested (Kannada is Sarvam-M's lone win, by a fraction of a point) and sweeps all 11 languages against Sarvam-30B. The widest gap against Sarvam-M shows up exactly where India's own models should have the edge: Odia, by 4.72 points (80.83% vs. 76.11%).

  4. The "winner" doesn't win everywhere: Sarvam-M still beats Qwen3.6-27B at Law & Governance and Social Sciences.

    Qwen3.6-27B's overall lead hides a split scoreboard by subject. It crushes Sarvam-M 24B in Science, its widest domain gap in either direction (93.38% vs. 86.05%, +7.33 points) — but flip to Law & Governance or Social Sciences and Sarvam-M comes out on top instead, by 2.99 and 2.21 points. Those are arguably the two domains where local legal and civic context matters most, and they're exactly where India's own model holds its ground.

  5. DeepSeek delivers solid accuracy at a fraction of a frontier model's price.

    DeepSeek V4-Flash scored 79.5% while costing only ₹2.45 per 1,000 answers, real metered spend — a small fraction of what any closed frontier model in our comparison costs (₹31-190/1K), even though it also trails all of them on accuracy (they run 87-94%).

  6. Indian languages still trail English — Odia most of all.

    Every model scored higher in English than across the 10 Indian languages tested. Hindi comes closest to English; Odia shows the largest gap — more than twice Hindi's.

Also worth noting

  • Newer generations of the same-sized models show clear gains (typically 4–8 points).
  • Health & Medicine is the strongest subject for most models examined so far — Qwen3.8-Max is the one exception, where Science edges ahead instead. Arts & Humanities and Social Sciences trail the top domain by 6–16 points in every model, no exceptions yet — the widest and most consistent gap in the test.

findings

01

Score it wrong, and your best model looks like your worst.

Sarvam-M 24B, gpt-oss-20b, and Qwen3.6-27B each use a chat template built around an internal reasoning step, which the loglikelihood protocol can't score correctly regardless of whether reasoning is switched on. We rescored all three generatively instead, since that's the protocol that actually fits how they're meant to run. Sarvam-M 24B and gpt-oss-20b are scored with thinking allowed — letting each model actually reason before answering — landing at #3 (82.44%) and #4 (72.22%). We scored Qwen3.6-27B both ways: thinking deliberately turned off (#5, 67.00%) and thinking on (#2 overall, #1 open-weight, 83.84%) — same checkpoint, +16.84 points from one config flag, the best open-weight result in the whole roster. We report all of these under the generative protocol, since that's the only one that scores them meaningfully. Two other rows in the roster are also generative, for unrelated reasons — Mistral Small 3.1 24B (its tokenizer can't survive the loglikelihood harness's context/continuation round-trip) and Sarvam-30B (never run under loglikelihood at all) — the remaining ten rows use the standard loglikelihood protocol. (DeepSeek V4-Flash and Qwen3.8-Max, the two models in our roster accessed only via API, are covered separately in the frontier & API comparison below.)

02

3× the parameters, half the gains: Gemma 3's law of diminishing returns.

We measured +13.4 points average going 4B→12B (3× params), but only +7.5 going 12B→27B (2.25× params) — a clear diminishing-returns curve, though the second step is also a smaller relative jump in parameters, so the two gains aren't directly comparable per unit of scale. Odia trails English by ~21-23 points at every size we tested; scaling doesn't touch the low-resource gap.

03

Gemma 4 beat Gemma 3 in 10 languages out of 11. Tamil didn't get the memo.

We found a real architecture/training gain at matched 12B scale, not a scale effect. Nine languages gained 4.6–7.2 points; English gained less than that band (+2.56) and Tamil barely moved at all (−0.02) — both below the other nine. Tamil's result is effectively flat (−0.02), an anomaly worth investigating.

04

Shrink the model, and Odia pays 7× what Bengali does.

We measured a 2.47-point overall loss from quantizing Gemma 3 12B to INT4, landing unevenly across languages: Odia −4.42, Tamil −3.55, Kannada −3.29, versus Bengali −0.63 — Odia lost more than 7× what Bengali did. That's not simply a low-resource-vs-English story, though: English itself lost 3.22 points, in the same range as Tamil and Kannada, and the overall English-Indic gap actually narrowed slightly under quantization (14.24 → 13.43). The unevenness is real; which languages take the bigger hit doesn't cleanly track resource level.

05

Half the parameters, the same score as last year's flagship.

We measured 56.41%, against the paper's published 56.90% — inside a point, at less than half the size. Not a clean per-language sweep (Gemma 3 12B trails on English/Gujarati/Hindi by 1-3 points) — an average-level story, not a uniform one.

06

Every model we've broken down by domain shares the same Achilles' heel.

Health & Medicine is the strongest domain for every model we've broken down by domain so far, except Qwen3.8-Max, where Science edges ahead instead. The weakest is always Arts & Humanities or Social Sciences (6-16 point gap), no exception yet. Model family, size, none of it changes which domains break a model.

07

17B active parameters go toe-to-toe with 27B dense — and nearly win.

We measured Llama 4 Scout (MoE, 17B active of 109B total) at 63.06% — within 0.9 points of Gemma 3 27B's 63.95%, a fully dense model with 10B more active parameters per token. Consistent with active-parameter efficiency, not just a scale story — though with only one MoE model in the roster, we can't isolate routing itself from other architecture/training differences between these two model families.

08

Qwen's generational leap is nearly double Gemma's, at near-identical size.

We measured Qwen3-VL 8B Instruct at 50.33% against Qwen2.5 7B Instruct's 41.81% — a generational jump at near-identical parameter count, structurally similar to the matched-size Gemma 3→4 comparison above but nearly double its magnitude (+8.52 vs. +4.72). Consistent with a generational improvement rather than pure scale — though unlike the matched-architecture Gemma comparison, Qwen3-VL's architecture also differs from Qwen2.5's, so we can't fully separate training from architecture here.

09

Sarvam's newer, bigger model just lost to its own smaller sibling.

Sarvam-30B (~32B total parameters, ~2.4B active per token, mixture-of-experts) scores 72.99% overall — 9.45 points behind Sarvam's own dense 24B model, Sarvam-M (82.44%). More total parameters didn't help: Sarvam-M activates all 24B of its parameters per token, versus Sarvam-30B's ~2.4B, and Sarvam-30B's number also comes from a compressed (Q4_K_M) checkpoint rather than full precision — either factor could be doing some of the work.

MILU accuracy vs cost per 1,000 answered questions

We normalized to a per-question unit cost rather than total run cost, so a huge dense model and a tiny sparse one are comparable on the same axis. Every model on this chart is open-weight and hosted locally, so cost is GPU-hours × an illustrative reference rate ($1.80/hr, one on-demand A100 80GB PCIe instance) — not a real bill, since we ran these on a box that's privately provisioned and shared. We converted to ₹ at ₹95.29/$1, then divided by thousands of successfully-answered questions across all 11 languages (79,608 total per model). For loglikelihood models, we measured this GPU-hour figure as the wall-clock time of the full run, which already includes all 4 per-option forward passes per question — dividing that real total by the number of questions (not by 4× as many requests) is what makes a loglikelihood row and a generative row comparable on the same axis. The ₹ conversion itself is a real rate; the underlying $ cost basis is still not — treat every ₹ figure here as directional, not an invoice. The two models we accessed only via API (Qwen3.8-Max, DeepSeek V4-Flash) use real metered spend instead of this reference rate, so we break them out separately in the frontier & API comparison above rather than mixing cost bases on one axis.

We numbered each point by league-table rank (matches the # column in the table further down this page) and connected its model name with a thin leader line, routed into open space rather than sitting on top of the dot. A few points (#09-#11) sit close together in both accuracy and cost — that's genuine proximity in the underlying data; follow the leader line or hover for the exact number. The x-axis extends to ₹200 to fit the roster's two priciest points, Qwen3.6-27B, thinking on (#01, 83.84%) at ₹161.35/1K and Sarvam-30B (#03, 72.99%) at ₹165.03/1K — near-coincidentally close in cost (both are 27B+-class models served via llama.cpp across the full 79,608-item set at similar GPU-hour totals) despite an 11-point accuracy gap between them. A log-scale x-axis handles that stretch without compressing the cheaper points on the left.

₹1 ₹10 ₹100 ₹200 20% 30% 40% 50% 60% 70% 80% 90% cost per 1,000 answered questions (₹, log scale) MILU overall accuracy, item-weighted (%) #01 Qwen3.6-27B, thinking on: 83.84% at ₹161.35/1K answered questions #02 Sarvam-M 24B: 82.44% at ₹41.85/1K answered questions #03 Sarvam-30B: 72.99% at ₹165.03/1K answered questions #04 gpt-oss-20b: 72.22% at ₹19.71/1K answered questions #05 Qwen3.6-27B (thinking off): 67.00% at ₹8.16/1K answered questions #06 Mistral Small 3.1 24B: 64.04% at ₹1.59/1K answered questions #07 Gemma 3 27B: 63.95% at ₹28.55/1K answered questions #08 Llama 4 Scout: 63.06% at ₹47.40/1K answered questions #09 Gemma 4 12B: 61.12% at ₹15.47/1K answered questions #10 Gemma 3 12B: 56.41% at ₹14.46/1K answered questions #11 Gemma 3 12B INT4: 53.94% at ₹13.96/1K answered questions #12 Qwen3-VL 8B: 50.33% at ₹18.79/1K answered questions #13 Phi-4 14B: 47.55% at ₹31.43/1K answered questions #14 Gemma 3 4B: 43.03% at ₹5.41/1K answered questions #15 Qwen2.5 7B: 41.81% at ₹17.07/1K answered questions #16 Sarvam-1 2B: 28.63% at ₹3.17/1K answered questions 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 Qwen3.6-27B, thinking on Sarvam-M 24B Sarvam-30B gpt-oss-20b Qwen3.6-27B (off) Mistral Small 3.1 24B Gemma 3 27B Llama 4 Scout Gemma 4 12B Gemma 3 12B Gemma 3 12B INT4 Qwen3-VL 8B Phi-4 14B Gemma 3 4B Qwen2.5 7B Sarvam-1 2B
local/GPU, open-weight (reference rate)
Mistral Small 3.1 24B is the cheapest point on this chart at ₹1.59/1K, but trails well behind on accuracy (64.04%). At the other end, Sarvam-30B (₹165.03/1K) and Qwen3.6-27B with thinking on (₹161.35/1K, our best open-weight model by accuracy at 83.84%) are the two priciest points on the whole chart — both local/GPU checkpoints, just at 27B+ scale with reasoning enabled. We left the two API-accessed models we ran ourselves (Qwen3.8-Max, DeepSeek V4-Flash) and the closed frontier models (Claude, GPT, Gemini) off this chart — see the dedicated frontier & API comparison above.

frontier & API model comparison

This section mixes two different kinds of sourcing, kept clearly separate in the table below. Eleven closed frontier models come from three system cards' own MILU comparison tables — Claude Opus 4.7, Claude Sonnet 5, and Claude Sonnet 4.6's own cards, plus every model each of those cards benchmarked itself against. We did not run any of these eleven — no local checkpoint exists for any of them, and API access wasn't in scope for this phase. We took their accuracy verbatim from each system card's own MILU comparison table — self-reported for that card's own Claude model, and run by that card's authors (Anthropic) for every non-Anthropic model shown alongside it, not self-reported by Google or OpenAI. We estimated their cost ourselves — not a bill — per the model's own tokenizer family: tiktoken's o200k_base (Anthropic and OpenAI models — exact for OpenAI, an approximation for Claude) or a local Gemma tokenizer (Google models — Gemini and Gemma share the same SentencePiece tokenizer family, a materially closer proxy than tiktoken, especially on Indic scripts). Qwen3.8-Max and DeepSeek V4-Flash are different — the two models in our own roster we accessed only via API, with no local checkpoint to host ourselves. Unlike the eleven above, we ran both of these ourselves, on the full 79,608-item MILU test set, with real metered $ spend, not an estimate — see the league table note for why they're grouped here with the API-accessed frontier models rather than the open-weight roster. The scatter chart below plots only the eleven vendor-cited frontier models, on their own scale, since that cluster (87-94% accuracy) sits well above our own roster; Qwen3.8-Max and DeepSeek appear in the table beneath it instead.

₹50 ₹100 ₹150 ₹200 ₹0 85% 87% 89% 91% 93% 95% est. ₹ per 1,000 answered questions (linear scale) MILU accuracy, as reported in system cards (%) Gemini 3.1 Pro: 93.6% (Claude Opus 4.7 system card, Table 8.12.2.B; dynamic thinking, default high level). Est. ₹31.78/1K -- Gemma-tokenizer proxy × published Gemini 3.1 Pro API pricing ($2.00/$12.00 per 1M tokens, ≤200K context). Gemini 3.1 Pro Gemini 3 Pro: 93.2% (Claude Sonnet 4.6 system card, Table 2.19.2.A; reasoning at default high effort). Est. ₹31.78/1K -- Gemma-tokenizer proxy × published Gemini 3 Pro API pricing ($2.00/$12.00 per 1M tokens, ≤200K context). Gemini 3 Pro GPT-5.4: 90.6% (Claude Opus 4.7 system card, Table 8.12.2.B; reasoning effort set to high). Est. ₹40.62/1K -- tiktoken o200k_base (exact for this model family) × published GPT-5.4 API pricing ($2.50/$15.00 per 1M tokens, ≤200K context). GPT-5.4 Claude Sonnet 4.6: 89.6% (Sonnet 4.6 system card, Table 2.19.2.A; adaptive thinking, max effort). Est. ₹47.03/1K -- tiktoken o200k_base proxy (not Claude's real tokenizer) × published pricing ($3/$15 per 1M tokens). Claude Sonnet 4.6 Claude Sonnet 5: 89.3% (Sonnet 5 system card, Figure 8.13.2.A; adaptive thinking, max effort, avg of 5 trials). Est. ₹47.03/1K at standard pricing ($3/$15 per 1M); ₹31.36/1K at introductory pricing ($2/$10 per 1M, in effect through 2026-08-31). tiktoken o200k_base proxy, not Claude's real tokenizer. Claude Sonnet 5 Claude Sonnet 4.5: 87.6% (Claude Sonnet 4.6 system card, Table 2.19.2.A; no adaptive thinking support, max thinking budget 1,024 tokens). Est. ₹47.03/1K -- tiktoken o200k_base proxy × published pricing ($3/$15 per 1M tokens, assumed same tier as Sonnet 4.6/5). Claude Sonnet 4.5 Claude Opus 4.8: 91.1% (Claude Sonnet 5 system card, Figure 8.13.2.A; adaptive thinking, max effort, avg of 5 trials). Est. ₹78.39/1K -- tiktoken o200k_base proxy × published pricing ($5/$25 per 1M tokens). Claude Opus 4.8 Claude Opus 4.7: 89.9% (Opus 4.7 system card, Table 8.12.2.B; adaptive thinking enabled). Est. ₹78.39/1K -- tiktoken o200k_base proxy × published pricing ($5/$25 per 1M tokens). Claude Opus 4.7 Claude Opus 4.6: 89.6% (Claude Sonnet 4.6 system card, Table 2.19.2.A; adaptive thinking, max effort). Est. ₹78.39/1K -- tiktoken o200k_base proxy × published pricing ($5/$25 per 1M tokens). Claude Opus 4.6 Claude Mythos Preview: 92.7% (Claude Sonnet 5 system card, Figure 8.13.2.A; adaptive thinking, max effort, avg of 5 trials). Est. ₹156.78/1K -- tiktoken o200k_base proxy × pricing assumed same tier as Claude Mythos 5/Fable 5 ($10/$50 per 1M tokens). Claude Mythos Preview GPT-5.2 Pro: 89.2% (Claude Sonnet 4.6 system card, Table 2.19.2.A; reasoning at default medium effort). Est. ₹182.63/1K -- tiktoken o200k_base (exact for this model family) × published GPT-5.2 Pro API pricing ($10.50/$84.00 per 1M tokens). GPT-5.2 Pro
Anthropic (tiktoken proxy) OpenAI (tiktoken, exact) Google (Gemma tokenizer proxy)
modelMILU accuracyest. ₹/1K answerssource
Gemini 3.1 Pro93.6%₹31.78Opus 4.7 card ↗, Table 8.12.2.B
Gemini 3 Pro93.2%₹31.78Sonnet 4.6 card ↗, Table 2.19.2.A
Claude Mythos Preview92.7%₹156.78Sonnet 5 card ↗, Fig. 8.13.2.A
Claude Opus 4.891.1%₹78.39Sonnet 5 card ↗, Fig. 8.13.2.A
GPT-5.490.6%₹40.62Opus 4.7 card ↗, Table 8.12.2.B
Claude Opus 4.789.9%₹78.39Opus 4.7 card ↗, Table 8.12.2.B
Claude Opus 4.689.6%₹78.39Sonnet 4.6 card ↗, Table 2.19.2.A
Claude Sonnet 4.689.6%₹47.03Sonnet 4.6 card ↗, Table 2.19.2.A
Claude Sonnet 589.3%₹47.03 (₹31.36 intro, through 2026-08-31)Sonnet 5 card ↗, Fig. 8.13.2.A
GPT-5.2 Pro89.2%₹182.63Sonnet 4.6 card ↗, Table 2.19.2.A
Claude Sonnet 4.587.6%₹47.03Sonnet 4.6 card ↗, Table 2.19.2.A
Qwen3.8-Max89.67%₹187.45we ran this — real spend
DeepSeek V4-Flash79.51%₹2.45we ran this — real spend
The eleven vendor-cited rows above report accuracy as an equal-weighted average across English + 10 Indic languages, exactly as each system card reports it — not item-count-weighted the way our own league table is, and not run under our own protocol. We estimated their cost ourselves — not real spend — from total tokens for the full 79,608-item MILU test set under our own 0-shot JSON-answer protocol, priced at each model's own published per-token API rate and converted at the same ₹95.29/$1 rate we use everywhere else in this report. Every vendor's reported accuracy used its own reasoning/thinking mode at a high or maximum setting, which bills extra output tokens our 0-shot cost estimate does not include — treat the accuracy and the cost figure in each of those eleven rows as two separate claims from two different protocols, not a matched cost-per-accuracy pair. Qwen3.8-Max and DeepSeek V4-Flash are the opposite on both counts: their accuracy is our own item-weighted measurement across the full 79,608 questions under our own generative protocol (same methodology as the league table), and their cost is real, metered API spend, not an estimate.

league table

This table covers 16 rows from the 15 distinct open-weight models in our roster — every one of them a checkpoint we can download and re-run ourselves; Qwen3.6-27B appears twice (#01 thinking on, #05 thinking off), a genuine ablation of the same checkpoint, not two different models. The two models we accessed only via API (Qwen3.8-Max, DeepSeek V4-Flash) are broken out separately in the frontier & API comparison above, alongside the closed frontier models — see that section for why. Six rows below are scored generative, 0-shot; ten are scored loglikelihood, 5-shot (the paper's own protocol) — see finding #01 above for why. The two protocols measure different things (open-ended generation vs. argmax over 4 fixed options), so ranks across that split are descriptive, not a strictly controlled comparison — read within-protocol comparisons with more confidence than across-protocol ones.

#model · pinned checkpoint in configsprotocoloverallteluguhindi₹ / 1K answers
01Qwen3.6-27B, thinking on opengenerative, 0-shot (with thinking)83.8483.5584.77₹161.35
02Sarvam-M 24B open · indiagenerative, 0-shot (with thinking)82.4481.4683.70₹41.85
03Sarvam-30B open · indiagenerative, 0-shot (with thinking)72.9971.2973.85₹165.03
04gpt-oss-20b opengenerative, 0-shot (with thinking)72.2270.7872.31₹19.71
05Qwen3.6-27B, thinking off opengenerative, 0-shot (thinking off)67.0058.7867.95₹8.16
06Mistral Small 3.1 24B opengenerative, 0-shot64.0459.6169.76₹1.59
07Gemma 3 27B openloglikelihood, 5-shot63.9562.1965.73₹28.55
08Llama 4 Scout 17B/109B openloglikelihood, 5-shot63.0659.8664.67₹47.40
09Gemma 4 12B openloglikelihood, 5-shot61.1257.7463.10₹15.47
10Gemma 3 12B openloglikelihood, 5-shot56.4153.1158.26₹14.46
11Gemma 3 12B INT4 openloglikelihood, 5-shot53.9450.3756.32₹13.96
12Qwen3-VL 8B Instruct openloglikelihood, 5-shot50.3346.1050.52₹18.79
13Phi-4 14B openloglikelihood, 5-shot47.5538.9450.50₹31.43
14Gemma 3 4B openloglikelihood, 5-shot43.0339.9545.52₹5.41
15Qwen2.5 7B Instruct openloglikelihood, 5-shot41.8134.8342.30₹17.07
16Sarvam-1 2B open · indialoglikelihood, 5-shot28.6328.4229.10₹3.17
₹/1K answers = cost per 1,000 successfully-answered questions across all 11 languages, GPU-hours × $1.80/hr reference rate, converted at an illustrative ₹95.29/$1 (not a real bill — see methodology). Every figure here is production-run cost only, excluding pre-production calibration spend. Qwen3.6-27B with thinking on (#01) and Sarvam-30B (#03) are the two standouts on cost, both 27B+-class checkpoints served via llama.cpp across all 11 languages independently rather than batched — Sarvam-30B's ₹165.03 is mostly two malformed-item recovery passes on top of the initial generation run (and excludes a separate one-time ~₹450 GPU spend from calibration work); Qwen3.6-27B's ₹161.35 is a single clean run with no retry passes needed. Qwen3.6-8B-Instruct wasn't available when we ran this, so we substituted Qwen3-VL-8B-Instruct.

breakdowns

This table shows accuracy on our four foregrounded product languages — Telugu, Hindi, Marathi, Kannada — side by side across nine selected models (all 17 in our roster completed full 11-language runs; these nine are shown here for space). For the per-domain (subject-area) pattern instead of per-language, see "the domain gap" section below.

modelteluguhindimarathikannada
Qwen3.8-Max89.0291.6788.0190.32
Qwen3.6-27B, thinking on83.5584.7782.1485.19
Sarvam-M 24B81.4683.7080.0085.58
Gemma 3 27B62.1965.7361.1962.48
Llama 4 Scout59.8664.6760.3867.84
Gemma 4 12B57.7463.1059.3360.75
Gemma 3 12B53.1158.2652.9555.44
Qwen3-VL 8B Instruct46.1050.5246.1747.56
Qwen2.5 7B Instruct34.8342.3037.1936.19
Sarvam-1 2B28.4229.1028.6727.43
Hindi leads for seven of the ten models shown, plausibly as the highest-resource of the four in most of these models' pretraining mix. Kannada leads for the other three — Llama 4 Scout (67.84%), Qwen3.6-27B with thinking on (85.19%), and Sarvam-M 24B (85.58%, also its single best language of all 11, not just of these four) — the opposite of the pattern the other seven models show, where Hindi is the strongest of the four.

the english-indic gap

We compare English against the 10-language Indic average, per model, across all 17 of our models — and, since every model shows some gap, which specific languages close it and which widen it.

modelenglishindic avggap
Sarvam-1 2B30.5428.132.41
Qwen3.8-Max91.6688.722.94
Qwen3.6-27B, thinking on86.0982.923.17
Sarvam-M 24B84.6881.483.21
gpt-oss-20b76.6471.045.60
Sarvam-30B77.3971.615.78
DeepSeek V4-Flash83.8077.945.86
Llama 4 Scout70.5060.819.69
Gemma 3 27B72.5061.3311.17
Gemma 4 12B69.9958.4911.50
Qwen3.6-27B, thinking off76.9964.8112.18
Gemma 3 12B INT464.2150.7813.43
Gemma 3 4B53.8239.7414.08
Gemma 3 12B67.4353.1814.24
Mistral Small 3.1 24B76.0959.7516.33
Qwen3-VL 8B64.7046.8617.84
Qwen2.5 7B60.5237.3423.19
Phi-4 14B66.5942.6423.95
The gap generally widens with lower overall accuracy, with two clusters bucking that trend: Sarvam-1 (weak everywhere, so little room to fall further), and a group of six stronger models — Qwen3.8-Max, Qwen3.6-27B with thinking on, Sarvam-M 24B, gpt-oss-20b, Sarvam-30B, and DeepSeek V4-Flash — all scoring 72%+ overall with a gap held under 6 points, well below what their neighbors on the accuracy scale show. Qwen3.8-Max has the smallest gap of any model above Sarvam-1 (2.94), consistent with it also being the highest-accuracy model in the roster; Qwen3.6-27B with thinking on is a close second-smallest (3.17) despite its thinking-off sibling showing a much wider gap (12.18) at the same checkpoint. Mistral is a notable case in the other direction: its overall gap (16.33) isn't unusual, but it's driven almost entirely by one language — we measured a 43.18-point drop from English on Odia alone (76.09% → 32.91%), by far the single largest per-language gap we saw for any model in our roster; every other language for Mistral trails English by a much more ordinary 6-19 points. Phi-4 14B is the largest model-level gap we measured (23.95, edging out Qwen2.5 7B's 23.19) despite a mid-table overall accuracy (47.55%) — its Indic scores trail English by a roughly even 21-28 points across every language, not one outlier language doing the work the way Mistral's Odia does.
languageavg gap from english
Hindi7.14
Bengali8.50
Kannada9.34
Gujarati9.95
Marathi11.37
Punjabi11.47
Telugu12.25
Malayalam12.98
Tamil13.50
Odia17.27
Averaged across all 17 models above, we found Hindi closes the English gap the most (7.14 points) — Odia widens it the most (17.27 points), more than 2.4× Hindi's gap. Every other language falls in between, in a fairly smooth ladder — consistent with Hindi being the highest-resource Indic language in most models' pretraining mix and Odia the lowest, across this specific set of models, though Mistral's own Odia figure is enough of an outlier that we'd treat it as a model-specific data point too, not purely a cross-model resource story.

the domain gap

Gemma 3 27B
73
66
61
59
Gemma 3 4B
54
44
40
41
Llama 4 Scout
73
65
60
57
Qwen3-VL 8B Instruct
61
54
47
44
We show four of the models we've broken down by domain so far, for space — domain breakdowns aren't yet computed for every model in the roster. Health & Medicine is the strongest domain for every model shown here. The weakest is Social Sciences for three of the four (Gemma 3 27B, Llama 4 Scout, Qwen3-VL 8B) — Gemma 3 4B is the exception, where Arts & Humanities (40.09) edges out Social Sciences (40.52) by less than half a point, close enough that we show both bars rather than asserting a single answer. The gap (strongest domain minus weakest) ranges 14.1-16.3 points across these four; the fuller 6-16 point range holds across our full analyzed set, and doesn't cleanly track model size or overall accuracy either way.

failure gallery

"Social Sciences" weakness is partly a logic-puzzle problem, not a knowledge gap. A large share of misses under it come from MILU's subject: Sociology bucket, which is mostly encoded family-relation logic puzzles:

$ Llama 4 Scout, Bengali — "If X is Y's sister; Y's mother is Z; Z's father is W; Z's husband is V; who is V's brother to Y?" — correct: "uncle." model said: "father."
$ Qwen3-VL 8B Instruct, Hindi — "If A*B means A is B's mother, A+B means A is B's sister, A%B means A is B's daughter — what does G+H*I%J mean?" — correct: "G is J's wife's sister." model said: "G is J's sister" (missed one hop).
$ Llama 4 Scout, Bengali (Sports and Recreation, plain factual miss for contrast) — "Which country hosted the 2017 Asian Women's Boxing Championship where India's Mary Kom won gold?" — correct: Vietnam. model said: China.

methodology

$ protocol    →  we used the MILU authors' lm-eval-harness task config, unmodified · 5-shot · temp 0 for loglikelihood rows; for generative rows (marked) we used 0-shot + JSON-answer extraction, MILU's own described API-scoring protocol
$ pinned     →  we pinned every local/open-weight model to an exact checkpoint revision, so results stay reproducible as new versions ship. DeepSeek V4-Flash (API-hosted, no local checkpoint) has no version pin available from the provider — its row reflects whatever the API served on our run date
$ cost basis  →  for API models we used real metered $ spend. For local models we used GPU-hours × $1.80/hr illustrative reference rate (one on-demand A100 80GB) — not a real bill, since we run this box privately and don't get billed per model. For our frontier reference models (11 closed models from Claude/OpenAI/Google system cards, see dedicated chart above) we didn't run anything at all — we estimated cost as a token-count against each model's own tokenizer family (tiktoken for Anthropic/OpenAI, a local Gemma tokenizer for Google) × published per-token pricing, not a bill of any kind; accuracy there is each vendor's own published number, not ours

credit & links

MILU is the work of AI4Bharat and IBM Research, published at NAACL 2025 (Verma et al., arXiv:2411.02538). They built the dataset, designed the task, and ran the original evaluation — none of this exists without that work, and we're adopters here, not authors.

What we added on top: we ran 17 models the original paper doesn't cover, including several 2026-era releases; we put a real cost-per-rupee lens next to raw accuracy, so a result means something on a budget, not just on a leaderboard; and we chased down whatever findings fell out of running both at once. If you cite one thing from this page, cite their paper first — this report is a downstream view of their benchmark, not a replacement for it.

We last updated this in Aug 2026, after rerunning Qwen3.6-27B with thinking mode switched on (a genuine ablation of a model already in the roster) — it took the #1 open-weight spot from Sarvam-M 24B. All 17 models, full 11-language MILU runs, 16 league-table rows.