We scored 21 off-the-shelf language models on the full Indic-KCC-Agri-Advisory-Benchmark — 500 Kisan Call Centre crop-advisory questions in 11 Indian languages, 5,500 graded answers per model, judged 1–5 on correctness, naturalness, groundedness and safety. deepseek-v4-flash leads at 4.29 correctness, hosted and undisclosed in size; among models we can actually run ourselves, gemma-3-27b-it and qwen3.6-27b tie exactly at 3.26, then split apart by axis — qwen3.6-27b posts the single highest safety score in the roster (4.86) while gemma-3-27b-it stays ahead on groundedness. Five models never clear 1.3 correctness, effectively unusable for this task. This survey excludes every checkpoint this project fine-tuned itself.
We ran the full Indic-KCC-Agri-Advisory-Benchmark — 500 Kisan Call Centre crop-advisory questions, translated into 11 Indian languages, 5,500 questions per model — against 21 off-the-shelf language models: open-weight and one hosted API model, none of them tuned on this data. We use an LLM-as-judge protocol — Qwen/Qwen3.6-35B-A3B-FP8 grades every answer, 1–5, on correctness, naturalness, groundedness and safety, independently.
deepseek-v4-flash leads the roster at 4.29 correctness — hosted, params undisclosed, and ahead of every other model on naturalness and groundedness too. Below it, gemma-3-27b-it and qwen3.6-27b tie exactly at 3.26: the best score among models this project can actually run itself. Below that, scores fall off steadily by size and training down to a floor cluster of five models that never clear 1.3 correctness. Safety barely moves across the whole roster (4.29–4.86) and doesn't track correctness at all — it's a weak differentiator here, not a meaningful signal of which model is "better." The English-vs-Indic gap is covered in detail below.
gemma-3-27b-it and qwen3.6-27b both land at exactly 3.26 correctness — but qwen3.6-27b posts 4.86 safety, the single highest score of any model in the roster, while gemma-3-27b-it leads on groundedness (4.16 vs 4.01) and naturalness (3.78 vs 3.75). Whichever one you'd call "better" depends entirely on which axis you weight.
gpt-oss-20b is a sparse mixture-of-experts model, 3.6B active out of 21B total parameters. At 1.99 correctness it beats every dense 7–8B instruct model we tested — llama3.1-8b-instruct (1.75), Qwen3-VL-8B-Instruct (1.73), krutrim-1-7b (1.56), navarasa-2.0-7b (1.38) — while running a fraction of the active compute per token.
Averaged across all 21 models, Odia trails English by 0.96 points of correctness, the widest gap of any of the 10 Indic languages (Hindi's is 0.34, the narrowest). mistral-small-3.1-24b takes that further: its Odia correctness collapses to 1.12 and safety to 3.26, against a 2.24–3.44 correctness range and 4.71–4.83 safety range everywhere else it was tested — despite a full 1.00 parse rate, so the failure isn't a broken pipeline.
Every model in the roster scores between 4.29 and 4.86 on safety, a 0.57-point band, versus a 3.21-point spread on correctness (1.08 to 4.29). Base safety alignment looks like it carries over almost unconditionally, regardless of how well a model actually answers the underlying agronomy question — safety score alone tells you almost nothing about answer quality here.
20 of the 21 models score higher in English than their own 10-language Indic average, some by more than a point and a half (Qwen__Qwen2.5-7B-Instruct: +1.77; Qwen__Qwen3-VL-8B-Instruct: +1.75). google__gemma-3-4b-it is the one exception: its English correctness (1.19) is actually 0.54 points below its own Indic average (1.73) — the only negative English-Indic gap in the whole roster.
Every one of the 21 models scores between 4.29 and 4.86 on safety — a band just 0.57 points wide. Correctness across the same roster spans 1.08 to 4.29, a 3.21-point range, nearly 6× wider. qwen3.6-27b holds the top safety score (4.86) without holding the top correctness score; kisanslm-gguf holds the bottom safety score (4.29) while still clearing several larger models on correctness. Base safety alignment looks like it survives almost regardless of how good the underlying agronomic answer actually is — on this benchmark, a safety number alone tells you very little about answer quality.
gpt-oss-20b is a sparse mixture-of-experts model — 3.6B active out of 21B total parameters per token. At 1.99 correctness it clears every dense 7–8B instruct model we tested: llama3.1-8b-instruct (1.75), Qwen3-VL-8B-Instruct (1.73), krutrim-1-7b (1.56) and navarasa-2.0-7b (1.38) all trail it, despite each running several times the active compute per token. It still sits well below the 24B+ tier (gemma-3-27b-it and qwen3.6-27b at 3.26), so this is active-parameter efficiency within its own size class, not a case for MoE closing the gap to much larger dense models.
Averaged across all 21 models, Odia trails English by 0.96 points of correctness — the widest gap of the 10 Indic languages tested, against Hindi's 0.34-point gap at the narrow end. mistral-small-3.1-24b shows what that can look like at the extreme: its Odia correctness collapses to 1.12 and safety to 3.26, against a 2.24–3.44 correctness range and 4.71–4.83 safety range everywhere else it answered — with a full 1.00 parse rate on that same Odia run, so the judge scored real output, not a parsing failure.
gemma-3-27b-it and qwen3.6-27b — both 27B, both open-weight — land on the exact same correctness score, 3.26. Break it down further and the tie dissolves: qwen3.6-27b posts 4.86 safety, the single highest score of any model in the whole roster, while gemma-3-27b-it leads on groundedness (4.16 vs 4.01) and naturalness (3.78 vs 3.75). Whichever one you'd call the better open-weight model for this task depends entirely on which axis matters most for your use case.
airavata-7b (1.08), openhathi-7b (1.16), sarvam-1-2b (1.23), param-1-2.9b (1.24) and qwen25vl-7b-base (1.28) all sit under 1.3 correctness — effectively unusable for crop-advisory answers as-is. It's not purely a small-model story: qwen25vl-7b-base is a 7B-class model (8.29B total, including its vision encoder), the same class as several mid-table models. Parse rates for this floor cluster stay reasonably high (0.95–0.98), so these are genuinely poor answers the judge could read cleanly, not malformed output dragging the score down.
deepseek-v4-flash is the only hosted, params-undisclosed model we tested, and it leads the entire roster at 4.29 correctness — 1.03 points ahead of the next-best score (3.26, tied between gemma-3-27b-it and qwen3.6-27b). It also leads on naturalness (4.38) and groundedness (4.67), the highest of any model on both axes. Its safety score (4.60) isn't the roster's highest, though — that distinction still goes to qwen3.6-27b (4.86).
All 21 models, sorted by correctness. Every row is an LLM-judge Stage 2 result: 500 questions × 11 languages, scored 1–5 on each axis (parse OK is a 0–1 share). Rows flagged anomaly are called out in the notes below the table.
| # | model · params | correctness | naturalness | groundedness | safety | parse ok |
|---|---|---|---|---|---|---|
| 01 | deepseek-v4-flash undisclosed (hosted) | 4.29 | 4.38 | 4.67 | 4.60 | 0.98 |
| 02 | gemma-3-27b-it 27B | 3.26 | 3.78 | 4.16 | 4.76 | 0.98 |
| 03 | qwen3.6-27b 27B | 3.26 | 3.75 | 4.01 | 4.86 | 0.99 |
| 04 | mistral-small-3.1-24b anomaly 24B | 2.54 | 2.66 | 3.33 | 4.63 | 0.99 |
| 05 | gemma3-12b-int4 12B (int4 QAT, base google/gemma-3-12b-it) | 2.40 | 3.10 | 3.43 | 4.74 | 0.98 |
| 06 | google · gemma-3-12b-it 12B | 2.30 | 3.22 | 3.44 | 4.84 | 0.99 |
| 07 | microsoft · phi-4 14B | 2.21 | 2.83 | 2.80 | 4.52 | 0.98 |
| 08 | gpt-oss-20b 21B total / 3.6B active (MoE) | 1.99 | 2.44 | 2.32 | 4.54 | 0.98 |
| 09 | llama3.1-8b-instruct 8B | 1.75 | 2.56 | 2.38 | 4.77 | 0.99 |
| 10 | Qwen · Qwen3-VL-8B-Instruct 8B | 1.73 | 2.20 | 2.15 | 4.80 | 0.98 |
| 11 | google · gemma-3-4b-it 4B | 1.68 | 1.98 | 2.20 | 4.76 | 0.98 |
| 12 | krutrim-1-7b 7B | 1.56 | 2.04 | 1.77 | 4.82 | 0.99 |
| 13 | sarvam-30b 30B | 1.51 | 1.39 | 2.00 | 4.64 | 0.96 |
| 14 | kisanslm-gguf 2B (Qwen3.5-2B base, Q4_K_M quant) | 1.47 | 1.76 | 1.58 | 4.29 | 0.98 |
| 15 | navarasa-2.0-7b 7B (Gemma-7B base) | 1.38 | 2.23 | 1.90 | 4.79 | 0.99 |
| 16 | Qwen · Qwen2.5-7B-Instruct 7.6B | 1.32 | 1.73 | 1.56 | 4.73 | 0.98 |
| 17 | qwen25vl-7b-base 8.29B (unsloth/Qwen2.5-VL-7B-Instruct) | 1.28 | 1.79 | 1.67 | 4.84 | 0.98 |
| 18 | param-1-2.9b 2.9B | 1.24 | 1.58 | 1.61 | 4.82 | 0.98 |
| 19 | sarvam-1-2b 2B | 1.23 | 1.33 | 1.63 | 4.72 | 0.95 |
| 20 | openhathi-7b 7B | 1.16 | 1.34 | 1.37 | 4.81 | 0.98 |
| 21 | airavata-7b 7B (maitreyaz/Airavata-8bit, OpenHathi-7B base) | 1.08 | 1.76 | 1.34 | 4.75 | 0.97 |
benchmark/results/indic_agri/comparison.md, Stage 2 aggregation. Excludes every checkpoint this project fine-tuned itself (the original LoRA checkpoint and every epoch/step checkpoint from both fine-tune sweeps) and sarvam-m-thinking. google__gemma-3-12b-it was judged twice under two different generation-token budgets; only its higher-scoring run is shown here, so this table has one row per model.We compare English correctness against the 10-language Indic average, per model, across all 21 models — and which specific languages close that gap versus widen it.
| model | english | indic avg | gap |
|---|---|---|---|
| google · gemma-3-4b-it | 1.19 | 1.73 | -0.54 |
| sarvam-30b | 1.54 | 1.51 | +0.03 |
| krutrim-1-7b | 1.69 | 1.54 | +0.15 |
| deepseek-v4-flash | 4.43 | 4.28 | +0.15 |
| sarvam-1-2b | 1.47 | 1.21 | +0.26 |
| gpt-oss-20b | 2.25 | 1.96 | +0.29 |
| gemma-3-27b-it | 3.66 | 3.22 | +0.44 |
| google · gemma-3-12b-it | 2.76 | 2.25 | +0.51 |
| navarasa-2.0-7b | 1.87 | 1.33 | +0.54 |
| airavata-7b | 1.61 | 1.03 | +0.58 |
| gemma3-12b-int4 | 3.08 | 2.33 | +0.75 |
| openhathi-7b | 1.89 | 1.09 | +0.80 |
| microsoft · phi-4 | 2.96 | 2.13 | +0.83 |
| llama3.1-8b-instruct | 2.53 | 1.67 | +0.86 |
| qwen3.6-27b | 4.08 | 3.18 | +0.90 |
| kisanslm-gguf | 2.35 | 1.38 | +0.97 |
| mistral-small-3.1-24b | 3.44 | 2.45 | +0.99 |
| param-1-2.9b | 2.15 | 1.15 | +1.00 |
| qwen25vl-7b-base | 2.60 | 1.15 | +1.46 |
| Qwen · Qwen3-VL-8B-Instruct | 3.32 | 1.56 | +1.75 |
| Qwen · Qwen2.5-7B-Instruct | 2.93 | 1.16 | +1.77 |
| language | avg gap from english |
|---|---|
| Hindi | 0.34 |
| Marathi | 0.53 |
| Bengali | 0.59 |
| Tamil | 0.64 |
| Gujarati | 0.66 |
| Punjabi | 0.67 |
| Telugu | 0.78 |
| Kannada | 0.80 |
| Malayalam | 0.93 |
| Odia | 0.96 |
Correctness, naturalness, groundedness and safety for a spread of models across the roster — the top scorer, the tied open-weight pair, the mid-table, and the floor.
A single-language, full-axis collapse in mistral-small-3.1-24b. Every axis for this model craters on Odia specifically — correctness 1.12, naturalness 1.15, groundedness 1.18, safety 3.26 — against a 2.24–3.44 correctness range and 4.71–4.83 safety range in every other language it answered. Parse rate on that same Odia run is a full 1.00, so the judge scored the answers it saw; the drop looks like a language-specific refusal or formatting failure in the model's own output, not a broken evaluation pipeline.
Qwen/Qwen3.6-35B-A3B-FP8, reference-grounded, scoring correctness, naturalness, groundedness and safety independently on a fixed 1–5 scaleGitHub: benchmark & eval harness →GitHub: model configs & leaderboard →Hugging Face dataset