Vendor model cards and published benchmarks are not a reliable basis for selecting a speech-to-text model. This report presents results from direct, reproducible evaluation — identical audio and identical scoring methodology throughout — across 16 open-weights speech-to-text models against AI4Bharat's IndicVoices dataset, covering as many of its 22 languages as each model supports. Four of the sixteen models produce a measurable but meaningless WER score: they translate, romanize, or reformat the input instead of transcribing it, a failure mode no WER number alone reveals.
Parameter count does not predict transcription accuracy — model architecture does.
The best-performing model in this benchmark is a 600M-parameter CTC conformer (IndicConformer, 22.02% WER, pooled across all 22 languages). A 430M-parameter model (SraVaani 1.0, 22.91% WER) trails by half a point despite covering only 19 of the 22 languages. Both outperform every autoregressive or LLM-decoder model evaluated, including models an order of magnitude larger. Parameter count was not the determining factor; whether the model was purpose-built for transcription was.
This distinction has practical consequences. Four additional models — all autoregressive, each marketed with some form of multilingual speech or speech-understanding capability — loaded and executed without error, producing fluent output that was not, in fact, a transcription of the input audio. A WER score computed against such output is technically valid but substantively meaningless, since it scores the wrong task. Section 2 documents all four cases, with source audio and model output.
Parameter count versus outcome, for the five models with confirmed published parameter counts:
| Model | Params | Outcome |
|---|---|---|
| SraVaani 1.0 | 430M | 22.91% WER — effectively ties the leader |
| IndicConformer | 600M | 22.02% WER — best overall |
| Qwen3-ASR-1.7B | 1.7B | 22.50% WER — still in the same band |
| Voxtral Mini 3B | 3B | 33.74% WER — a real number, well behind |
| Krutrim Dhwani-1 | 7.35B | translates to English rather than transcribing; the resulting WER score is not meaningful |
Increasing parameter count from 430M to 1.7B yields negligible improvement. At 7.35B, the model fails the task altogether: Dhwani-1's own repository ships an explicit transcription prompt template, yet the model translates the input regardless.
All four models loaded, executed, and produced fluent, confident output. Each failure was confirmed at the raw model-API level rather than in this study's inference code, and is therefore attributable to the model, not the evaluation harness.
Dhwani-1 represents a fifth instance of the translation failure mode and is documented in Section 1 rather than repeated here.
The IndicVoices paper (arXiv 2403.01926) reports its own IndicASR — a 130M-parameter conformer trained only on IndicVoices — at 25.65% macro WER, in its Table 7. We ran the newer, public 600M checkpoint of the same lineage (IndicConformer) and scored it the same way.
| Source | Model | Macro WER |
|---|---|---|
| Paper, Table 7 | IndicASR, 130M | 25.65% |
| This study | IndicConformer, 600M | 22.82% |
An improvement of 2.83 points on 20 of 22 languages, consistent with a larger, later-generation checkpoint. Two languages moved in the opposite direction — Kannada (+4.8pt) and Odia (+5.4pt) — the only anomaly in this comparison, not yet root-caused. Exact reproduction of the paper's results is not possible in either direction: the paper does not disclose its normalization scheme, decoding mode, or which transcript level (verbatim or standardized) it scores against.
SeamlessM4T, Dhwani-1, and Whisper Turbo each list the languages on which they subsequently failed. A model card describes training coverage, not behavior under a transcription instruction. The only reliable verification is direct evaluation against real audio.
A 430M-parameter CTC model built for a single task (SraVaani) matches a 600M sibling and outperforms every general-purpose speech model evaluated, several of them 5–15x larger. For the specific task of transcribing Indian-language speech, purpose-built models consistently outperform general-purpose alternatives.
zero-stt-hinglish recognizes most spoken content correctly yet scores 80.5% WER, because IndicVoices' conversational recordings contain frequent spoken phone numbers and one-time passcodes, which the model renders as digits rather than the spelled-out words the reference transcripts use. A model's output format must match the evaluation convention, independent of whether the underlying recognition is accurate.
Vakyansh (2021–22 vintage) is the only model in this benchmark with checkpoints for Sanskrit, Dogri, and Maithili — and its error rate runs roughly 3x that of the leading models. For these languages, the practical choice may be between limited-accuracy coverage and no coverage at all; no accurate option currently exists.
Micro WER pools every language's errors and words before dividing, so higher-resource languages contribute proportionally more. CER (character error rate) provides a complementary signal — see Section 6 for a worked example of why the two metrics can diverge. Bold indicates the two models within half a point of the leading score.
| Model | Org | Languages | Micro WER | CER |
|---|---|---|---|---|
| IndicConformer | AI4Bharat | 22 | 22.02% | 8.16% |
| SraVaani 1.0 | ARTPARK-IISc | 19 | 22.91% | 8.59% |
| Qwen3-ASR-1.7B | Alibaba | 1 (Hindi) | 22.50% | 11.02% |
| w2v-bert-punjabi | Independent | 1 (Punjabi) | 23.51% | 8.77% |
| Nemotron 3.5 ASR | NVIDIA | 1 (Hindi) | 24.16% | 12.87% |
| Shrutam-2 | BharatGen | 12 | 27.74% | 14.65% |
| Voxtral Mini 3B | Mistral | 1 (Hindi) | 33.74% | 20.33% |
| IndicWav2Vec | AI4Bharat | 4 | 48.00% | 20.29% |
| Vakyansh | Open-Speech-EkStep | 16 | 63.70% | 29.36% |
| Meta MMS | Meta | 15 | 65.26% | 31.28% |
| Whisper Large-v3-Turbo | OpenAI | 14 | 138.38% · script drift | 102.38% |
| SeamlessM4T v2 | Meta | multiple | not comparable · translates | — |
| Krutrim Dhwani-1 | Krutrim AI Labs | multiple | not comparable · translates | — |
| zero-stt-hinglish | Shunya Labs | 1 (Hindi) | 80.54% · format mismatch | 69.14% |
→ scroll for the full table on small screens
Three commercial, pay-per-use APIs were evaluated using the same audio and scoring methodology as every open-weights model above, rather than relying on vendor-published figures. Results are presented in two parts: Hindi (full split, all three vendors), followed by six additional languages (all three vendors again, since all three support this language set). All costs are shown in Indian rupees; Soniox and Deepgram bill in USD and are converted here at ₹96/US$1 (September 2026).
Hindi, full split, all three vendors:
| Rank | Model | WER | CER | Cost |
|---|---|---|---|---|
| 1 | IndicConformer (600M) | 14.40% | 6.73% | open-weights, local GPU |
| 2 | SraVaani 1.0 (430M) | 14.50% | 6.83% | open-weights, local GPU |
| 3 | Sarvam saaras:v4 | 14.68% | 7.43% | ~₹248–332 · full split |
| 4 | Shrutam-2 | 18.20% | 10.82% | open-weights, local GPU |
| 5 | Soniox stt-async-v5 | 19.42% | 10.37% | ₹85 · full split, 5,530 utt |
| 6 | Qwen3-ASR-1.7B | 22.50% | 11.02% | open-weights, local GPU |
| 7 | Nemotron 3.5 ASR | 24.16% | 12.87% | open-weights, local GPU |
| 8 | Deepgram nova-3 | 25.60% | 17.47% | ₹220 · full split, 5,530 utt |
| 9 | Voxtral Mini 3B | 33.74% | 20.33% | open-weights, local GPU |
Sarvam ranks third, effectively tied with the two leading open-weights models. Soniox ranks fifth, ahead of every Hindi-only purpose-built open-weights model evaluated (Qwen3-ASR, Nemotron, Voxtral). Deepgram ranks last among the nine. All three commercial results reflect full-scale evaluation runs.
Additional Languages: Soniox vs. Deepgram vs. Sarvam (Bengali, Gujarati, Kannada, Marathi, Tamil, Telugu)
Evaluated using identical audio and scoring methodology against all three vendors' full valid splits. Odia was initially included in this comparison but excluded, as neither Soniox nor Deepgram supports it: Soniox returns a 400 "Invalid language hint" error for the Odia language code, confirmed directly against the live API and consistent with its absence from Soniox's published language list; Deepgram returns the same error code (see coverage table below).
| Language | Soniox WER | Soniox CER | Deepgram WER | Deepgram CER | Sarvam WER | Sarvam CER |
|---|---|---|---|---|---|---|
| Bengali | 19.58% | 9.42% | 26.62% | 15.75% | 12.86% | 5.71% |
| Gujarati | 21.30% | 9.20% | 20.34% | 8.06% | 16.73% | 6.24% |
| Marathi | 23.73% | 10.24% | 32.42% | 17.91% | 14.94% | 6.03% |
| Telugu | 36.23% | 13.64% | 28.12% | 10.62% | 25.16% | 8.70% |
| Kannada | 49.58% | 17.10% | 43.49% | 15.30% | 33.29% | 10.54% |
| Tamil* | 49.91% | 16.79% | 35.02% | 12.17% | 32.10% | 9.97% |
| Micro | 34.10% | 13.41% | 31.39% | 13.41% | 22.85% | 8.21% |
| Cost | ₹378 | — | ₹975 | — | ₹1,102 | — |
→ scroll to see all three vendors on small screens
*For Soniox, the Tamil sample size is n=5,275 rather than 5,276: one utterance failed due to a request timeout unrelated to the Odia issue, and was not re-attempted given its negligible effect on the aggregate result.
Sarvam wins every language in this set, several by a wide margin — its micro WER (22.85%) approaches the pooled scores of the open-weights leaders (IndicConformer 22.02%, SraVaani 22.91%, each across a wider language set), and it costs less than Deepgram while beating it on every single language. Between Soniox and Deepgram alone, neither is uniformly superior: Soniox performs better on Bengali and Marathi; Deepgram performs better on Telugu, Kannada, and particularly Tamil. Kannada and Tamil remain Soniox's weakest results in this set by a substantial margin — both Dravidian-script languages, both near 50% WER, while its results on Indo-Aryan-script languages stay in the low-to-mid twenties, a pattern consistent with a genuine script or language-family effect rather than noise. Combined with its Hindi result (Section 6, above), Sarvam is now the strongest commercial API in this benchmark on every language tested.
Deepgram Language Coverage: 11 of 22 IndicVoices Languages
Confirmed by direct testing against the live API — each unsupported language code returns a 400 error — since Deepgram's documentation does not enumerate IndicVoices-specific coverage.
| Supported | Not supported |
|---|---|
| Assamese, Bengali, Gujarati, Hindi, Kannada, Marathi, Nepali, Punjabi, Tamil, Telugu, Urdu | Bodo, Dogri, Konkani, Kashmiri, Maithili, Malayalam, Manipuri, Odia, Sanskrit, Santali, Sindhi |
Malayalam and Odia are notable omissions given their relatively high resource availability; overall coverage caps at exactly half of the IndicVoices language set. The remaining unsupported languages form the same lower-resource cluster that presents difficulty across vendors generally.
On Hindi, Deepgram's CER-to-WER ratio is 0.68 (17.47% CER against 25.60% WER); every other model in the Hindi comparison falls between 0.47 (IndicConformer) and 0.60 (Voxtral). An elevated ratio typically indicates that incorrect words are not close phonetic substitutions, or that automated formatting — Deepgram's smart_format feature governs punctuation and numeral rendering — diverges from the reference transcript's conventions, the same category of issue documented for zero-stt-hinglish in Section 2, though less severe here. Unlike other findings in this report, this has not yet been confirmed through direct transcript inspection.
Results for the six models with broad multi-language coverage. A dash indicates no available checkpoint for that language. Vakyansh is the only model in this set with any coverage of Sanskrit, Dogri, or Maithili.
| Language | IndicConformer | SraVaani | Shrutam-2 | MMS | IndicWav2Vec | Vakyansh |
|---|---|---|---|---|---|---|
| Assamese | 16.6 | 20.7 | 26.1 | 56.0 | — | 67.8 |
| Bengali | 13.5 | 13.3 | 18.4 | 46.8 | 46.5 | 54.3 |
| Bodo | 22.4 | 25.7 | — | — | — | — |
| Dogri | 30.6 | 32.7 | — | — | — | 88.9 |
| Gujarati | 17.4 | 19.6 | 25.0 | 45.8 | 41.0 | 47.0 |
| Hindi | 14.4 | 14.5 | 18.2 | 42.5 | 37.4 | 34.9 |
| Kannada | 35.1 | 34.3 | 43.8 | 80.0 | — | 78.6 |
| Konkani | 26.4 | 28.3 | — | — | — | — |
| Kashmiri | 33.0 | — | — | — | — | — |
| Maithili | 29.7 | 30.9 | — | 81.2 | — | 91.3 |
| Malayalam | 36.2 | 37.3 | 45.3 | 79.0 | — | 77.5 |
| Manipuri | 18.7 | 21.7 | — | — | — | — |
| Marathi | 14.9 | 15.2 | 20.5 | 56.5 | — | 79.6 |
| Nepali | 16.2 | 17.0 | — | 66.6 | — | 69.5 |
| Odia | 28.8 | 29.3 | 39.3 | 68.7 | 72.1 | 78.3 |
| Punjabi | 11.0 | 12.2 | 18.4 | 45.3 | — | 56.3 |
| Sanskrit | 19.6 | 21.5 | — | — | — | 69.7 |
| Santali | 28.8 | 30.9 | — | 104.8 | — | — |
| Sindhi | 22.3 | 23.2 | — | 101.4 | — | — |
| Tamil | 30.8 | — | 42.4 | 84.4 | — | 66.2 |
| Telugu | 26.3 | 26.6 | 30.7 | 68.1 | — | 57.9 |
| Urdu | 8.9 | — | 13.8 | — | — | 37.8 |
→ scroll to see all 22 languages
One representative utterance per language, selected as the utterance with the broadest model coverage in this benchmark. Ground truth, model output, and WER/CER are drawn directly from the same prediction data underlying every figure in this report. This interactive view applies the same transcript-verification method used to identify the failure modes documented in Section 2.
| Model | Transcript | WER | CER |
|---|---|---|---|
| IndicConformerAI4Bharat | দু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার | 25.0%wer | 2.2%cer |
| SraVaani 1.0ARTPARK-IISc | দু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার | 25.0%wer | 2.2%cer |
| Soniox stt-async-v5Soniox | ২৫০০০০, ৩৪০০০০, ২০৩০, ২০০৫, ১৫০০০০। | 100.0%wer | 95.6%cer |
| Shrutam-2BharatGen | দু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার | 25.0%wer | 2.2%cer |
| Deepgram nova-3Deepgram | 2 লক্ষ পঞ্চাশ হাজার, 3 লক্ষ চল্লিশ হাজার, 2৩০, 2 পাঁচ, 1 লক্ষ পঞ্চাশ হাজার | 37.5%wer | 30.0%cer |
| IndicWav2VecAI4Bharat | ২লক্ ৫০ হাজার ৩ন লকখ৪০ হাজার ২ুহাজার৩০ ২হাজার ৫চ ১ক লকষ৫০ হাজার | 81.2%wer | 46.7%cer |
| VakyanshOpen-Speech-EkStep | দু লক্ষ পঞ্চাস হাজার তিন লক্ষ চল্লিস হাজার দুই হাজার তিরিস দুই হাজার পাঁচ এক লক্ষ পঞ্চাস হাজার | 50.0%wer | 8.9%cer |
| Meta MMSMeta | 2লকষ পঞচাশ হাজার 3িন লক্ষ 4ি০,াজার 2,ার3 2,5াচ এ লক্ষ পঞচাশ হাজার | 75.0%wer | 37.8%cer |
Ground truth. IndicVoices provides two transcript levels per utterance: verbatim (phonetic, as spoken) and standardized (normalized orthography). This benchmark scores against the standardized level, consistent with the expectation that a transcription model should output normalized text rather than colloquial spelling.
Normalization. Text is normalized by stripping Unicode punctuation categories (covering sentence-final marks across all scripts, including the Devanagari danda ।) and casefolding Latin-script segments. Numeral expansion is not applied.
Metric. Word error rate and character error rate are computed using the jiwer library. Both micro-averaged (pooled) and macro-averaged (per-language mean) results are reported, as the IndicVoices paper does not specify which convention it uses. Character error rate counts whitespace as a character — jiwer's default behavior, computing Levenshtein distance over the full normalized string rather than word-delimited text. This is stated explicitly because it affects the reported value: see the Section 8 Telugu example, where CER is 16.67% excluding spaces and 20.00% (the value this report uses) including them.
Split. Only the train and valid splits of IndicVoices are publicly available; no held-out test split exists. All results in this report are computed against the full valid split.
Ground-truth transcripts and audio are from IndicVoices (AI4Bharat, CC BY 4.0), described in arXiv:2403.01926. The scoring methodology (stt_bench), every model integration, and every inference run reported here were developed independently for this study. No reported figure is drawn from a vendor model card or published benchmark.
All figures are computed on IndicVoices' public valid split; no held-out test split is available. Any reproduction of these results necessarily scores against data that models trained by third parties may have encountered during development. This cannot be audited for models this study did not train.
The percentages describing each failure mode in Section 2 (for example, "46% of Hindi utterances translated") are computed from the full scored run for that model. The specific audio/transcript examples shown were selected for clarity rather than sampled at random.
This benchmark evaluates transcription accuracy and documented failure modes; it does not evaluate latency or deployment readiness. A model with strong transcription accuracy here may still be unsuitable for a given deployment for other reasons. Commercial API costs reported in Section 6 reflect this study's actual spend at each provider's published pay-as-you-go rate.