MILU, re-run on the 2026 models.
summary
We ran the full 11-language MILU test set — 79,608 questions each — on 17 models available through mid-2026 (16 league-table rows: 15 open-weight models, with Qwen3.6-27B appearing twice, thinking on and off — the other 2 models were API-only and are covered separately below). We matched the scoring protocol to each model: loglikelihood 5-shot where the harness fits, generative 0-shot where it doesn't, and paired every result with a real or reference-derived cost per 1,000 answered questions.
Qwen3.6-27B with thinking on leads the open-weight league table at 83.84% overall, ahead of Sarvam-M 24B's 82.44%. Its own thinking-off configuration scores 16.84 points lower (67.00%), the widest swing from a single setting change in the roster. Two models in our roster — Qwen3.8-Max (89.67%, ₹187.45/1K) and DeepSeek V4-Flash (79.51%, ₹2.45/1K) — were accessed only via API with no local checkpoint, so we compare them separately against closed frontier models rather than ranking them alongside models we can actually re-host ourselves. Beyond the rankings, the data show that scoring protocol is the main confounder for reasoning models, matched-size generations still gain meaningfully (Gemma 3→4, Qwen2.5→Qwen3-VL), and the low-resource gap — especially Odia — persists across almost every model regardless of scale.
key insights
Alibaba's Qwen3.6-27B leads India's own best open model on our leaderboard.
Qwen3.6-27B, run with thinking enabled, tops our open-weight leaderboard at 83.84% overall, ahead of Sarvam-M 24B's 82.44% — India's strongest open model. The catch: it isn't cheap. At ₹161.35 per 1,000 answers, it costs nearly 4× what Sarvam-M does (₹41.85/1K) — a real trade-off, not a free win.
Two models tied within a point overall — but who's "ahead" flips by subject.
OpenAI's open-weight gpt-oss-20b trails Sarvam-30B by less than a point overall (72.22% vs 72.99%). Break it down by domain and the lead flips: gpt-oss-20b actually wins Science (+2.6 points) and Social Sciences (+1.4 points). Sarvam-30B pulls furthest ahead in Law & Governance, by 5.9 points — of every foreign-vs-Indic open-model pair we checked, this is the only one where the domain-level lead changes hands at all.
Alibaba's model beats India's own flagship in 9 of 10 Indian languages — and it isn't close in Odia.
Qwen3.6-27B doesn't just edge past Sarvam-M 24B on the overall score — it out-scores Sarvam-M in 9 of the 10 Indic languages we tested (Kannada is Sarvam-M's lone win, by a fraction of a point) and sweeps all 11 languages against Sarvam-30B. The widest gap against Sarvam-M shows up exactly where India's own models should have the edge: Odia, by 4.72 points (80.83% vs. 76.11%).
The "winner" doesn't win everywhere: Sarvam-M still beats Qwen3.6-27B at Law & Governance and Social Sciences.
Qwen3.6-27B's overall lead hides a split scoreboard by subject. It crushes Sarvam-M 24B in Science, its widest domain gap in either direction (93.38% vs. 86.05%, +7.33 points) — but flip to Law & Governance or Social Sciences and Sarvam-M comes out on top instead, by 2.99 and 2.21 points. Those are arguably the two domains where local legal and civic context matters most, and they're exactly where India's own model holds its ground.
DeepSeek delivers solid accuracy at a fraction of a frontier model's price.
DeepSeek V4-Flash scored 79.5% while costing only ₹2.45 per 1,000 answers, real metered spend — a small fraction of what any closed frontier model in our comparison costs (₹31-190/1K), even though it also trails all of them on accuracy (they run 87-94%).
Indian languages still trail English — Odia most of all.
Every model scored higher in English than across the 10 Indian languages tested. Hindi comes closest to English; Odia shows the largest gap — more than twice Hindi's.
Also worth noting
- Newer generations of the same-sized models show clear gains (typically 4–8 points).
- Health & Medicine is the strongest subject for most models examined so far — Qwen3.8-Max is the one exception, where Science edges ahead instead. Arts & Humanities and Social Sciences trail the top domain by 6–16 points in every model, no exceptions yet — the widest and most consistent gap in the test.
findings
Score it wrong, and your best model looks like your worst.
Sarvam-M 24B, gpt-oss-20b, and Qwen3.6-27B each use a chat template built around an internal reasoning step, which the loglikelihood protocol can't score correctly regardless of whether reasoning is switched on. We rescored all three generatively instead, since that's the protocol that actually fits how they're meant to run. Sarvam-M 24B and gpt-oss-20b are scored with thinking allowed — letting each model actually reason before answering — landing at #3 (82.44%) and #4 (72.22%). We scored Qwen3.6-27B both ways: thinking deliberately turned off (#5, 67.00%) and thinking on (#2 overall, #1 open-weight, 83.84%) — same checkpoint, +16.84 points from one config flag, the best open-weight result in the whole roster. We report all of these under the generative protocol, since that's the only one that scores them meaningfully. Two other rows in the roster are also generative, for unrelated reasons — Mistral Small 3.1 24B (its tokenizer can't survive the loglikelihood harness's context/continuation round-trip) and Sarvam-30B (never run under loglikelihood at all) — the remaining ten rows use the standard loglikelihood protocol. (DeepSeek V4-Flash and Qwen3.8-Max, the two models in our roster accessed only via API, are covered separately in the frontier & API comparison below.)
3× the parameters, half the gains: Gemma 3's law of diminishing returns.
We measured +13.4 points average going 4B→12B (3× params), but only +7.5 going 12B→27B (2.25× params) — a clear diminishing-returns curve, though the second step is also a smaller relative jump in parameters, so the two gains aren't directly comparable per unit of scale. Odia trails English by ~21-23 points at every size we tested; scaling doesn't touch the low-resource gap.
Gemma 4 beat Gemma 3 in 10 languages out of 11. Tamil didn't get the memo.
We found a real architecture/training gain at matched 12B scale, not a scale effect. Nine languages gained 4.6–7.2 points; English gained less than that band (+2.56) and Tamil barely moved at all (−0.02) — both below the other nine. Tamil's result is effectively flat (−0.02), an anomaly worth investigating.
Shrink the model, and Odia pays 7× what Bengali does.
We measured a 2.47-point overall loss from quantizing Gemma 3 12B to INT4, landing unevenly across languages: Odia −4.42, Tamil −3.55, Kannada −3.29, versus Bengali −0.63 — Odia lost more than 7× what Bengali did. That's not simply a low-resource-vs-English story, though: English itself lost 3.22 points, in the same range as Tamil and Kannada, and the overall English-Indic gap actually narrowed slightly under quantization (14.24 → 13.43). The unevenness is real; which languages take the bigger hit doesn't cleanly track resource level.
Half the parameters, the same score as last year's flagship.
We measured 56.41%, against the paper's published 56.90% — inside a point, at less than half the size. Not a clean per-language sweep (Gemma 3 12B trails on English/Gujarati/Hindi by 1-3 points) — an average-level story, not a uniform one.
Every model we've broken down by domain shares the same Achilles' heel.
Health & Medicine is the strongest domain for every model we've broken down by domain so far, except Qwen3.8-Max, where Science edges ahead instead. The weakest is always Arts & Humanities or Social Sciences (6-16 point gap), no exception yet. Model family, size, none of it changes which domains break a model.
17B active parameters go toe-to-toe with 27B dense — and nearly win.
We measured Llama 4 Scout (MoE, 17B active of 109B total) at 63.06% — within 0.9 points of Gemma 3 27B's 63.95%, a fully dense model with 10B more active parameters per token. Consistent with active-parameter efficiency, not just a scale story — though with only one MoE model in the roster, we can't isolate routing itself from other architecture/training differences between these two model families.
Qwen's generational leap is nearly double Gemma's, at near-identical size.
We measured Qwen3-VL 8B Instruct at 50.33% against Qwen2.5 7B Instruct's 41.81% — a generational jump at near-identical parameter count, structurally similar to the matched-size Gemma 3→4 comparison above but nearly double its magnitude (+8.52 vs. +4.72). Consistent with a generational improvement rather than pure scale — though unlike the matched-architecture Gemma comparison, Qwen3-VL's architecture also differs from Qwen2.5's, so we can't fully separate training from architecture here.
Sarvam's newer, bigger model just lost to its own smaller sibling.
Sarvam-30B (~32B total parameters, ~2.4B active per token, mixture-of-experts) scores 72.99% overall — 9.45 points behind Sarvam's own dense 24B model, Sarvam-M (82.44%). More total parameters didn't help: Sarvam-M activates all 24B of its parameters per token, versus Sarvam-30B's ~2.4B, and Sarvam-30B's number also comes from a compressed (Q4_K_M) checkpoint rather than full precision — either factor could be doing some of the work.
MILU accuracy vs cost per 1,000 answered questions
We normalized to a per-question unit cost rather than total run cost, so a huge dense model and a tiny sparse one are comparable on the same axis. Every model on this chart is open-weight and hosted locally, so cost is GPU-hours × an illustrative reference rate ($1.80/hr, one on-demand A100 80GB PCIe instance) — not a real bill, since we ran these on a box that's privately provisioned and shared. We converted to ₹ at ₹95.29/$1, then divided by thousands of successfully-answered questions across all 11 languages (79,608 total per model). For loglikelihood models, we measured this GPU-hour figure as the wall-clock time of the full run, which already includes all 4 per-option forward passes per question — dividing that real total by the number of questions (not by 4× as many requests) is what makes a loglikelihood row and a generative row comparable on the same axis. The ₹ conversion itself is a real rate; the underlying $ cost basis is still not — treat every ₹ figure here as directional, not an invoice. The two models we accessed only via API (Qwen3.8-Max, DeepSeek V4-Flash) use real metered spend instead of this reference rate, so we break them out separately in the frontier & API comparison above rather than mixing cost bases on one axis.
We numbered each point by league-table rank (matches the # column in the table further down this page) and connected its model name with a thin leader line, routed into open space rather than sitting on top of the dot. A few points (#09-#11) sit close together in both accuracy and cost — that's genuine proximity in the underlying data; follow the leader line or hover for the exact number. The x-axis extends to ₹200 to fit the roster's two priciest points, Qwen3.6-27B, thinking on (#01, 83.84%) at ₹161.35/1K and Sarvam-30B (#03, 72.99%) at ₹165.03/1K — near-coincidentally close in cost (both are 27B+-class models served via llama.cpp across the full 79,608-item set at similar GPU-hour totals) despite an 11-point accuracy gap between them. A log-scale x-axis handles that stretch without compressing the cheaper points on the left.
frontier & API model comparison
This section mixes two different kinds of sourcing, kept clearly separate in the table below. Eleven closed frontier models come from three system cards' own MILU comparison tables — Claude Opus 4.7, Claude Sonnet 5, and Claude Sonnet 4.6's own cards, plus every model each of those cards benchmarked itself against. We did not run any of these eleven — no local checkpoint exists for any of them, and API access wasn't in scope for this phase. We took their accuracy verbatim from each system card's own MILU comparison table — self-reported for that card's own Claude model, and run by that card's authors (Anthropic) for every non-Anthropic model shown alongside it, not self-reported by Google or OpenAI. We estimated their cost ourselves — not a bill — per the model's own tokenizer family: tiktoken's o200k_base (Anthropic and OpenAI models — exact for OpenAI, an approximation for Claude) or a local Gemma tokenizer (Google models — Gemini and Gemma share the same SentencePiece tokenizer family, a materially closer proxy than tiktoken, especially on Indic scripts). Qwen3.8-Max and DeepSeek V4-Flash are different — the two models in our own roster we accessed only via API, with no local checkpoint to host ourselves. Unlike the eleven above, we ran both of these ourselves, on the full 79,608-item MILU test set, with real metered $ spend, not an estimate — see the league table note for why they're grouped here with the API-accessed frontier models rather than the open-weight roster. The scatter chart below plots only the eleven vendor-cited frontier models, on their own scale, since that cluster (87-94% accuracy) sits well above our own roster; Qwen3.8-Max and DeepSeek appear in the table beneath it instead.
| model | MILU accuracy | est. ₹/1K answers | source |
|---|---|---|---|
| Gemini 3.1 Pro | 93.6% | ₹31.78 | Opus 4.7 card ↗, Table 8.12.2.B |
| Gemini 3 Pro | 93.2% | ₹31.78 | Sonnet 4.6 card ↗, Table 2.19.2.A |
| Claude Mythos Preview | 92.7% | ₹156.78 | Sonnet 5 card ↗, Fig. 8.13.2.A |
| Claude Opus 4.8 | 91.1% | ₹78.39 | Sonnet 5 card ↗, Fig. 8.13.2.A |
| GPT-5.4 | 90.6% | ₹40.62 | Opus 4.7 card ↗, Table 8.12.2.B |
| Claude Opus 4.7 | 89.9% | ₹78.39 | Opus 4.7 card ↗, Table 8.12.2.B |
| Claude Opus 4.6 | 89.6% | ₹78.39 | Sonnet 4.6 card ↗, Table 2.19.2.A |
| Claude Sonnet 4.6 | 89.6% | ₹47.03 | Sonnet 4.6 card ↗, Table 2.19.2.A |
| Claude Sonnet 5 | 89.3% | ₹47.03 (₹31.36 intro, through 2026-08-31) | Sonnet 5 card ↗, Fig. 8.13.2.A |
| GPT-5.2 Pro | 89.2% | ₹182.63 | Sonnet 4.6 card ↗, Table 2.19.2.A |
| Claude Sonnet 4.5 | 87.6% | ₹47.03 | Sonnet 4.6 card ↗, Table 2.19.2.A |
| Qwen3.8-Max | 89.67% | ₹187.45 | we ran this — real spend |
| DeepSeek V4-Flash | 79.51% | ₹2.45 | we ran this — real spend |
league table
This table covers 16 rows from the 15 distinct open-weight models in our roster — every one of them a checkpoint we can download and re-run ourselves; Qwen3.6-27B appears twice (#01 thinking on, #05 thinking off), a genuine ablation of the same checkpoint, not two different models. The two models we accessed only via API (Qwen3.8-Max, DeepSeek V4-Flash) are broken out separately in the frontier & API comparison above, alongside the closed frontier models — see that section for why. Six rows below are scored generative, 0-shot; ten are scored loglikelihood, 5-shot (the paper's own protocol) — see finding #01 above for why. The two protocols measure different things (open-ended generation vs. argmax over 4 fixed options), so ranks across that split are descriptive, not a strictly controlled comparison — read within-protocol comparisons with more confidence than across-protocol ones.
| # | model · pinned checkpoint in configs | protocol | overall | telugu | hindi | ₹ / 1K answers |
|---|---|---|---|---|---|---|
| 01 | Qwen3.6-27B, thinking on open | generative, 0-shot (with thinking) | 83.84 | 83.55 | 84.77 | ₹161.35 |
| 02 | Sarvam-M 24B open · india | generative, 0-shot (with thinking) | 82.44 | 81.46 | 83.70 | ₹41.85 |
| 03 | Sarvam-30B open · india | generative, 0-shot (with thinking) | 72.99 | 71.29 | 73.85 | ₹165.03 |
| 04 | gpt-oss-20b open | generative, 0-shot (with thinking) | 72.22 | 70.78 | 72.31 | ₹19.71 |
| 05 | Qwen3.6-27B, thinking off open | generative, 0-shot (thinking off) | 67.00 | 58.78 | 67.95 | ₹8.16 |
| 06 | Mistral Small 3.1 24B open | generative, 0-shot | 64.04 | 59.61 | 69.76 | ₹1.59 |
| 07 | Gemma 3 27B open | loglikelihood, 5-shot | 63.95 | 62.19 | 65.73 | ₹28.55 |
| 08 | Llama 4 Scout 17B/109B open | loglikelihood, 5-shot | 63.06 | 59.86 | 64.67 | ₹47.40 |
| 09 | Gemma 4 12B open | loglikelihood, 5-shot | 61.12 | 57.74 | 63.10 | ₹15.47 |
| 10 | Gemma 3 12B open | loglikelihood, 5-shot | 56.41 | 53.11 | 58.26 | ₹14.46 |
| 11 | Gemma 3 12B INT4 open | loglikelihood, 5-shot | 53.94 | 50.37 | 56.32 | ₹13.96 |
| 12 | Qwen3-VL 8B Instruct open | loglikelihood, 5-shot | 50.33 | 46.10 | 50.52 | ₹18.79 |
| 13 | Phi-4 14B open | loglikelihood, 5-shot | 47.55 | 38.94 | 50.50 | ₹31.43 |
| 14 | Gemma 3 4B open | loglikelihood, 5-shot | 43.03 | 39.95 | 45.52 | ₹5.41 |
| 15 | Qwen2.5 7B Instruct open | loglikelihood, 5-shot | 41.81 | 34.83 | 42.30 | ₹17.07 |
| 16 | Sarvam-1 2B open · india | loglikelihood, 5-shot | 28.63 | 28.42 | 29.10 | ₹3.17 |
breakdowns
This table shows accuracy on our four foregrounded product languages — Telugu, Hindi, Marathi, Kannada — side by side across nine selected models (all 17 in our roster completed full 11-language runs; these nine are shown here for space). For the per-domain (subject-area) pattern instead of per-language, see "the domain gap" section below.
| model | telugu | hindi | marathi | kannada |
|---|---|---|---|---|
| Qwen3.8-Max | 89.02 | 91.67 | 88.01 | 90.32 |
| Qwen3.6-27B, thinking on | 83.55 | 84.77 | 82.14 | 85.19 |
| Sarvam-M 24B | 81.46 | 83.70 | 80.00 | 85.58 |
| Gemma 3 27B | 62.19 | 65.73 | 61.19 | 62.48 |
| Llama 4 Scout | 59.86 | 64.67 | 60.38 | 67.84 |
| Gemma 4 12B | 57.74 | 63.10 | 59.33 | 60.75 |
| Gemma 3 12B | 53.11 | 58.26 | 52.95 | 55.44 |
| Qwen3-VL 8B Instruct | 46.10 | 50.52 | 46.17 | 47.56 |
| Qwen2.5 7B Instruct | 34.83 | 42.30 | 37.19 | 36.19 |
| Sarvam-1 2B | 28.42 | 29.10 | 28.67 | 27.43 |
the english-indic gap
We compare English against the 10-language Indic average, per model, across all 17 of our models — and, since every model shows some gap, which specific languages close it and which widen it.
| model | english | indic avg | gap |
|---|---|---|---|
| Sarvam-1 2B | 30.54 | 28.13 | 2.41 |
| Qwen3.8-Max | 91.66 | 88.72 | 2.94 |
| Qwen3.6-27B, thinking on | 86.09 | 82.92 | 3.17 |
| Sarvam-M 24B | 84.68 | 81.48 | 3.21 |
| gpt-oss-20b | 76.64 | 71.04 | 5.60 |
| Sarvam-30B | 77.39 | 71.61 | 5.78 |
| DeepSeek V4-Flash | 83.80 | 77.94 | 5.86 |
| Llama 4 Scout | 70.50 | 60.81 | 9.69 |
| Gemma 3 27B | 72.50 | 61.33 | 11.17 |
| Gemma 4 12B | 69.99 | 58.49 | 11.50 |
| Qwen3.6-27B, thinking off | 76.99 | 64.81 | 12.18 |
| Gemma 3 12B INT4 | 64.21 | 50.78 | 13.43 |
| Gemma 3 4B | 53.82 | 39.74 | 14.08 |
| Gemma 3 12B | 67.43 | 53.18 | 14.24 |
| Mistral Small 3.1 24B | 76.09 | 59.75 | 16.33 |
| Qwen3-VL 8B | 64.70 | 46.86 | 17.84 |
| Qwen2.5 7B | 60.52 | 37.34 | 23.19 |
| Phi-4 14B | 66.59 | 42.64 | 23.95 |
| language | avg gap from english |
|---|---|
| Hindi | 7.14 |
| Bengali | 8.50 |
| Kannada | 9.34 |
| Gujarati | 9.95 |
| Marathi | 11.37 |
| Punjabi | 11.47 |
| Telugu | 12.25 |
| Malayalam | 12.98 |
| Tamil | 13.50 |
| Odia | 17.27 |
the domain gap
failure gallery
"Social Sciences" weakness is partly a logic-puzzle problem, not a knowledge gap. A large share of misses under it come from MILU's subject: Sociology bucket, which is mostly encoded family-relation logic puzzles:
methodology
credit & links
What we added on top: we ran 17 models the original paper doesn't cover, including several 2026-era releases; we put a real cost-per-rupee lens next to raw accuracy, so a result means something on a budget, not just on a leaderboard; and we chased down whatever findings fell out of running both at once. If you cite one thing from this page, cite their paper first — this report is a downstream view of their benchmark, not a replacement for it.
We last updated this in Aug 2026, after rerunning Qwen3.6-27B with thinking mode switched on (a genuine ablation of a model already in the roster) — it took the #1 open-weight spot from Sarvam-M 24B. All 17 models, full 11-language MILU runs, 16 league-table rows.
