16 ASR MODELS EVALUATED · MEASURED WITH REAL INFERENCE, NOT SELF-REPORTED
speech-to-text benchmark
An Independent Benchmark of 16 Indic Speech-to-Text Models on IndicVoices report india

Vendor model cards and published benchmarks are not a reliable basis for selecting a speech-to-text model. This report presents results from direct, reproducible evaluation — identical audio and identical scoring methodology throughout — across 16 open-weights speech-to-text models against AI4Bharat's IndicVoices dataset, covering as many of its 22 languages as each model supports. Four of the sixteen models produce a measurable but meaningless WER score: they translate, romanize, or reformat the input instead of transcribing it, a failure mode no WER number alone reveals.

report · IndicVoices, Level 2 standardized transcripts · 357,000+ utterances scored · Sept 2026
speech-to-text benchmark

What It Actually Said.

16 open-weights ASR models · IndicVoices, 22 languages · word error rate measured on real audio, not self-reported figures

1. The Main Finding

Parameter count does not predict transcription accuracy — model architecture does.

The best-performing model in this benchmark is a 600M-parameter CTC conformer (IndicConformer, 22.02% WER, pooled across all 22 languages). A 430M-parameter model (SraVaani 1.0, 22.91% WER) trails by half a point despite covering only 19 of the 22 languages. Both outperform every autoregressive or LLM-decoder model evaluated, including models an order of magnitude larger. Parameter count was not the determining factor; whether the model was purpose-built for transcription was.

This distinction has practical consequences. Four additional models — all autoregressive, each marketed with some form of multilingual speech or speech-understanding capability — loaded and executed without error, producing fluent output that was not, in fact, a transcription of the input audio. A WER score computed against such output is technically valid but substantively meaningless, since it scores the wrong task. Section 2 documents all four cases, with source audio and model output.

Parameter count versus outcome, for the five models with confirmed published parameter counts:

Parameter count compared with observed outcome on IndicVoices audio.
ModelParamsOutcome
SraVaani 1.0430M22.91% WER — effectively ties the leader
IndicConformer600M22.02% WER — best overall
Qwen3-ASR-1.7B1.7B22.50% WER — still in the same band
Voxtral Mini 3B3B33.74% WER — a real number, well behind
Krutrim Dhwani-17.35Btranslates to English rather than transcribing; the resulting WER score is not meaningful

Increasing parameter count from 430M to 1.7B yields negligible improvement. At 7.35B, the model fails the task altogether: Dhwani-1's own repository ships an explicit transcription prompt template, yet the model translates the input regardless.

Best overall · micro WER
22.02%
IndicConformer, 600M · all 22 languages
Best per parameter
22.91%
SraVaani 1.0, 430M · 19 languages
Utterances scored
357K+
across every completed run
Models ruled out
6 / 16
translation, script drift, format mismatch, or unreliable output
IndicConformerAI4Bharat · 22 langs
22.02%
SraVaani 1.0ARTPARK-IISc · 19 langs
22.91%
Qwen3-ASR-1.7BAlibaba · Hindi only
22.50%
w2v-bert-punjabiIndependent · Punjabi only
23.51%
Nemotron 3.5 ASRNVIDIA · Hindi only
24.16%
Shrutam-2BharatGen · 12 langs
27.74%
Voxtral Mini 3BMistral · Hindi only
33.74%
IndicWav2VecAI4Bharat · 4 langs
48.00%
VakyanshOpen-Speech-EkStep · 16 langs
63.70%
Meta MMSMeta · 15 langs
65.26%
Every model producing a valid transcription, ranked by micro WER. Pooled across each model's own language set, so higher-resource languages carry proportionally more weight. The four models whose output does not constitute a genuine transcription — and whose WER scores are therefore not meaningful — are covered separately in Section 2.

2. Failure Modes Not Captured by WER Alone

All four models loaded, executed, and produced fluent, confident output. Each failure was confirmed at the raw model-API level rather than in this study's inference code, and is therefore attributable to the model, not the evaluation harness.

FAILS BY
Translating
SeamlessM4T v2 · Meta
46% of Hindi utterances were returned in English. "वहाँ पे धार्मिक भी है" → "There's religious stuff, too."
FAILS BY
Romanizing
Whisper Large-v3-Turbo · OpenAI
Assamese script was rendered in Latin characters on 52.7% of utterances, producing 138% WER.
FAILS BY
Reformatting
zero-stt-hinglish · Shunya Labs
Transcribed content is largely correct, but spoken digits are rendered as numerals ("3 7 5 4...") rather than the spelled-out words IndicVoices uses, producing 80.5% WER.
FAILS BY
Drifting
Shuka-1 · Sarvam (dropped, not scored)
An audio question-answering model rather than a transcriber. Intermittently failed to reference the provided audio, and in one case returned an unrelated safety refusal.

Dhwani-1 represents a fifth instance of the translation failure mode and is documented in Section 1 rather than repeated here.

3. Comparison Against the Original Paper's Baseline

The IndicVoices paper (arXiv 2403.01926) reports its own IndicASR — a 130M-parameter conformer trained only on IndicVoices — at 25.65% macro WER, in its Table 7. We ran the newer, public 600M checkpoint of the same lineage (IndicConformer) and scored it the same way.

SourceModelMacro WER
Paper, Table 7IndicASR, 130M25.65%
This studyIndicConformer, 600M22.82%

An improvement of 2.83 points on 20 of 22 languages, consistent with a larger, later-generation checkpoint. Two languages moved in the opposite direction — Kannada (+4.8pt) and Odia (+5.4pt) — the only anomaly in this comparison, not yet root-caused. Exact reproduction of the paper's results is not possible in either direction: the paper does not disclose its normalization scheme, decoding mode, or which transcript level (verbatim or standardized) it scores against.

4. Recommendations

  1. Listed language support does not guarantee transcription capability.

    SeamlessM4T, Dhwani-1, and Whisper Turbo each list the languages on which they subsequently failed. A model card describes training coverage, not behavior under a transcription instruction. The only reliable verification is direct evaluation against real audio.

  2. Purpose-built models outperform general-purpose models at comparable or larger scale.

    A 430M-parameter CTC model built for a single task (SraVaani) matches a 600M sibling and outperforms every general-purpose speech model evaluated, several of them 5–15x larger. For the specific task of transcribing Indian-language speech, purpose-built models consistently outperform general-purpose alternatives.

  3. Output formatting conventions materially affect measured accuracy.

    zero-stt-hinglish recognizes most spoken content correctly yet scores 80.5% WER, because IndicVoices' conversational recordings contain frequent spoken phone numbers and one-time passcodes, which the model renders as digits rather than the spelled-out words the reference transcripts use. A model's output format must match the evaluation convention, independent of whether the underlying recognition is accurate.

  4. Language coverage and accuracy trade off sharply at the low-resource end.

    Vakyansh (2021–22 vintage) is the only model in this benchmark with checkpoints for Sanskrit, Dogri, and Maithili — and its error rate runs roughly 3x that of the leading models. For these languages, the practical choice may be between limited-accuracy coverage and no coverage at all; no accurate option currently exists.

Two more models never produced a usable result

  • Meta Omnilingual ASR (300M CTC) — fast and otherwise clean, but intermittently outputs Urdu Perso-Arabic script for Hindi audio, a shared-vocabulary script-conditioning weakness.
  • Sarvam Shuka-1 — an audio question-answering model, not built for transcription; less consistent than outright mistranslation (see Section 2).

5. Full Results

Micro WER pools every language's errors and words before dividing, so higher-resource languages contribute proportionally more. CER (character error rate) provides a complementary signal — see Section 6 for a worked example of why the two metrics can diverge. Bold indicates the two models within half a point of the leading score.

ModelOrgLanguagesMicro WERCER
IndicConformerAI4Bharat2222.02%8.16%
SraVaani 1.0ARTPARK-IISc1922.91%8.59%
Qwen3-ASR-1.7BAlibaba1 (Hindi)22.50%11.02%
w2v-bert-punjabiIndependent1 (Punjabi)23.51%8.77%
Nemotron 3.5 ASRNVIDIA1 (Hindi)24.16%12.87%
Shrutam-2BharatGen1227.74%14.65%
Voxtral Mini 3BMistral1 (Hindi)33.74%20.33%
IndicWav2VecAI4Bharat448.00%20.29%
VakyanshOpen-Speech-EkStep1663.70%29.36%
Meta MMSMeta1565.26%31.28%
Whisper Large-v3-TurboOpenAI14138.38% · script drift102.38%
SeamlessM4T v2Metamultiplenot comparable · translates—
Krutrim Dhwani-1Krutrim AI Labsmultiplenot comparable · translates—
zero-stt-hinglishShunya Labs1 (Hindi)80.54% · format mismatch69.14%

→ scroll for the full table on small screens

6. Commercial API Comparison

Three commercial, pay-per-use APIs were evaluated using the same audio and scoring methodology as every open-weights model above, rather than relying on vendor-published figures. Results are presented in two parts: Hindi (full split, all three vendors), followed by six additional languages (all three vendors again, since all three support this language set). All costs are shown in Indian rupees; Soniox and Deepgram bill in USD and are converted here at ₹96/US$1 (September 2026).

Hindi, full split, all three vendors:

Hindi-only WER/CER and actual cost, ranked against every other model's own Hindi-specific score.
RankModelWERCERCost
1IndicConformer (600M)14.40%6.73%open-weights, local GPU
2SraVaani 1.0 (430M)14.50%6.83%open-weights, local GPU
3Sarvam saaras:v414.68%7.43%~₹248–332 · full split
4Shrutam-218.20%10.82%open-weights, local GPU
5Soniox stt-async-v519.42%10.37%₹85 · full split, 5,530 utt
6Qwen3-ASR-1.7B22.50%11.02%open-weights, local GPU
7Nemotron 3.5 ASR24.16%12.87%open-weights, local GPU
8Deepgram nova-325.60%17.47%₹220 · full split, 5,530 utt
9Voxtral Mini 3B33.74%20.33%open-weights, local GPU

Sarvam ranks third, effectively tied with the two leading open-weights models. Soniox ranks fifth, ahead of every Hindi-only purpose-built open-weights model evaluated (Qwen3-ASR, Nemotron, Voxtral). Deepgram ranks last among the nine. All three commercial results reflect full-scale evaluation runs.

Additional Languages: Soniox vs. Deepgram vs. Sarvam (Bengali, Gujarati, Kannada, Marathi, Tamil, Telugu)

Evaluated using identical audio and scoring methodology against all three vendors' full valid splits. Odia was initially included in this comparison but excluded, as neither Soniox nor Deepgram supports it: Soniox returns a 400 "Invalid language hint" error for the Odia language code, confirmed directly against the live API and consistent with its absence from Soniox's published language list; Deepgram returns the same error code (see coverage table below).

LanguageSoniox WERSoniox CERDeepgram WERDeepgram CERSarvam WERSarvam CER
Bengali19.58%9.42%26.62%15.75%12.86%5.71%
Gujarati21.30%9.20%20.34%8.06%16.73%6.24%
Marathi23.73%10.24%32.42%17.91%14.94%6.03%
Telugu36.23%13.64%28.12%10.62%25.16%8.70%
Kannada49.58%17.10%43.49%15.30%33.29%10.54%
Tamil*49.91%16.79%35.02%12.17%32.10%9.97%
Micro34.10%13.41%31.39%13.41%22.85%8.21%
Cost₹378—₹975—₹1,102—

→ scroll to see all three vendors on small screens

*For Soniox, the Tamil sample size is n=5,275 rather than 5,276: one utterance failed due to a request timeout unrelated to the Odia issue, and was not re-attempted given its negligible effect on the aggregate result.

Sarvam wins every language in this set, several by a wide margin — its micro WER (22.85%) approaches the pooled scores of the open-weights leaders (IndicConformer 22.02%, SraVaani 22.91%, each across a wider language set), and it costs less than Deepgram while beating it on every single language. Between Soniox and Deepgram alone, neither is uniformly superior: Soniox performs better on Bengali and Marathi; Deepgram performs better on Telugu, Kannada, and particularly Tamil. Kannada and Tamil remain Soniox's weakest results in this set by a substantial margin — both Dravidian-script languages, both near 50% WER, while its results on Indo-Aryan-script languages stay in the low-to-mid twenties, a pattern consistent with a genuine script or language-family effect rather than noise. Combined with its Hindi result (Section 6, above), Sarvam is now the strongest commercial API in this benchmark on every language tested.

Deepgram Language Coverage: 11 of 22 IndicVoices Languages

Confirmed by direct testing against the live API — each unsupported language code returns a 400 error — since Deepgram's documentation does not enumerate IndicVoices-specific coverage.

SupportedNot supported
Assamese, Bengali, Gujarati, Hindi, Kannada, Marathi, Nepali, Punjabi, Tamil, Telugu, UrduBodo, Dogri, Konkani, Kashmiri, Maithili, Malayalam, Manipuri, Odia, Sanskrit, Santali, Sindhi

Malayalam and Odia are notable omissions given their relatively high resource availability; overall coverage caps at exactly half of the IndicVoices language set. The remaining unsupported languages form the same lower-resource cluster that presents difficulty across vendors generally.

Deepgram's CER-to-WER Ratio Is an Outlier

On Hindi, Deepgram's CER-to-WER ratio is 0.68 (17.47% CER against 25.60% WER); every other model in the Hindi comparison falls between 0.47 (IndicConformer) and 0.60 (Voxtral). An elevated ratio typically indicates that incorrect words are not close phonetic substitutions, or that automated formatting — Deepgram's smart_format feature governs punctuation and numeral rendering — diverges from the reference transcript's conventions, the same category of issue documented for zero-stt-hinglish in Section 2, though less severe here. Unlike other findings in this report, this has not yet been confirmed through direct transcript inspection.

7. Per-Language Breakdown

Results for the six models with broad multi-language coverage. A dash indicates no available checkpoint for that language. Vakyansh is the only model in this set with any coverage of Sanskrit, Dogri, or Maithili.

LanguageIndicConformerSraVaaniShrutam-2MMSIndicWav2VecVakyansh
Assamese16.620.726.156.0—67.8
Bengali13.513.318.446.846.554.3
Bodo22.425.7————
Dogri30.632.7———88.9
Gujarati17.419.625.045.841.047.0
Hindi14.414.518.242.537.434.9
Kannada35.134.343.880.0—78.6
Konkani26.428.3————
Kashmiri33.0—————
Maithili29.730.9—81.2—91.3
Malayalam36.237.345.379.0—77.5
Manipuri18.721.7————
Marathi14.915.220.556.5—79.6
Nepali16.217.0—66.6—69.5
Odia28.829.339.368.772.178.3
Punjabi11.012.218.445.3—56.3
Sanskrit19.621.5———69.7
Santali28.830.9—104.8——
Sindhi22.323.2—101.4——
Tamil30.8—42.484.4—66.2
Telugu26.326.630.768.1—57.9
Urdu8.9—13.8——37.8

→ scroll to see all 22 languages

8. Transcript Examples

One representative utterance per language, selected as the utterance with the broadest model coverage in this benchmark. Ground truth, model output, and WER/CER are drawn directly from the same prediction data underlying every figure in this report. This interactive view applies the same transcript-verification method used to identify the failure modes documented in Section 2.

audio · Bengali
ground truth
দু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দুহাজার তিরিশ দুহাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার
ModelTranscriptWERCER
IndicConformerAI4Bharatদু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার25.0%wer2.2%cer
SraVaani 1.0ARTPARK-IIScদু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার25.0%wer2.2%cer
Soniox stt-async-v5Soniox২৫০০০০, ৩৪০০০০, ২০৩০, ২০০৫, ১৫০০০০।100.0%wer95.6%cer
Shrutam-2BharatGenদু লক্ষ পঞ্চাশ হাজার তিন লক্ষ চল্লিশ হাজার দু হাজার তিরিশ দু হাজার পাঁচ এক লক্ষ পঞ্চাশ হাজার25.0%wer2.2%cer
Deepgram nova-3Deepgram2 লক্ষ পঞ্চাশ হাজার, 3 লক্ষ চল্লিশ হাজার, 2৩০, 2 পাঁচ, 1 লক্ষ পঞ্চাশ হাজার37.5%wer30.0%cer
IndicWav2VecAI4Bharat২লক্ ৫০ হাজার ৩ন লকখ৪০ হাজার ২ুহাজার৩০ ২হাজার ৫চ ১ক লকষ৫০ হাজার81.2%wer46.7%cer
VakyanshOpen-Speech-EkStepদু লক্ষ পঞ্চাস হাজার তিন লক্ষ চল্লিস হাজার দুই হাজার তিরিস দুই হাজার পাঁচ এক লক্ষ পঞ্চাস হাজার50.0%wer8.9%cer
Meta MMSMeta2লকষ পঞচাশ হাজার 3িন লক্ষ 4ি০,াজার 2,ার3 2,5াচ এ লক্ষ পঞচাশ হাজার75.0%wer37.8%cer
IndicConformer, SraVaani, and Shrutam-2 all land at 25% WER here but just 2.2% CER — the text is nearly identical; the gap is a word-boundary spacing difference (।দু হাজার। vs ।দুহাজার।), not a real error. Soniox, on the exact same audio, scores 100% on *both* WER and CER: it transcribed the spoken number-words as digits (২৫০০০০ instead of দু লক্ষ পঞ্চাশ হাজার) — the same formatting mismatch that inflated zero-stt-hinglish's score in Section 2, independently caught here on a different vendor and a different language.

9. Methodology

Ground truth. IndicVoices provides two transcript levels per utterance: verbatim (phonetic, as spoken) and standardized (normalized orthography). This benchmark scores against the standardized level, consistent with the expectation that a transcription model should output normalized text rather than colloquial spelling.

Normalization. Text is normalized by stripping Unicode punctuation categories (covering sentence-final marks across all scripts, including the Devanagari danda ।) and casefolding Latin-script segments. Numeral expansion is not applied.

Metric. Word error rate and character error rate are computed using the jiwer library. Both micro-averaged (pooled) and macro-averaged (per-language mean) results are reported, as the IndicVoices paper does not specify which convention it uses. Character error rate counts whitespace as a character — jiwer's default behavior, computing Levenshtein distance over the full normalized string rather than word-delimited text. This is stated explicitly because it affects the reported value: see the Section 8 Telugu example, where CER is 16.67% excluding spaces and 20.00% (the value this report uses) including them.

Split. Only the train and valid splits of IndicVoices are publicly available; no held-out test split exists. All results in this report are computed against the full valid split.

Credit

Ground-truth transcripts and audio are from IndicVoices (AI4Bharat, CC BY 4.0), described in arXiv:2403.01926. The scoring methodology (stt_bench), every model integration, and every inference run reported here were developed independently for this study. No reported figure is drawn from a vendor model card or published benchmark.

10. Limitations

Scope and Limitations

All figures are computed on IndicVoices' public valid split; no held-out test split is available. Any reproduction of these results necessarily scores against data that models trained by third parties may have encountered during development. This cannot be audited for models this study did not train.

The percentages describing each failure mode in Section 2 (for example, "46% of Hindi utterances translated") are computed from the full scored run for that model. The specific audio/transcript examples shown were selected for clarity rather than sampled at random.

This benchmark evaluates transcription accuracy and documented failure modes; it does not evaluate latency or deployment readiness. A model with strong transcription accuracy here may still be unsuitable for a given deployment for other reasons. Commercial API costs reported in Section 6 reflect this study's actual spend at each provider's published pay-as-you-go rate.