Hundreds of millions of people in India type a mix of Hindi and English, in English letters — “mujhe kal ka schedule bhejo”. Many large AI models advertise Hindi or broader multilingual support; we could not find published figures for what romanized input costs them. We asked 16 AI models the same questions five different ways and measured the damage. The headline is not what the marketing implies: mixing languages is not the problem — dropping the Hindi script is.
Form 5 is a form many people use on a phone keyboard. It is also the one models handle worst, by a wide margin.
Writing in English letters costs more than mixing languages does — and the two costs compound.
Hinglish was the easiest of the four non-English forms for every model we could meaningfully rank — all 15 on general knowledge and all 13 on maths. Models kept 86.5% of their ability (95% CI 84.2–88.8) on general knowledge and 95.7% (94.9–96.6) on maths. “Kept” has a precise meaning, defined just before the results tables below. A Hinglish question gives a model two footholds: familiar English words, and familiar Hindi script. Either one substantially cushions the loss — but neither removes it, and mixing is not free: Hinglish still costs 13.5 points on general knowledge.
Which models are excluded, and why. 3 small models score so close to guessing that ranking language forms by them is meaningless — Sarvam-1 2B, Param-1 2.9B, OpenHathi 7B on maths, OpenHathi 7B on general knowledge. OpenHathi 7B solves about 1 maths problem in 16 in English, and Param-1 2.9B scores higher on Hinglish (9.21) than on English (6.18), which cannot be a real effect. Those rows are dimmed in the tables and excluded from every count and retention figure. Two of them do reverse the pattern on maths: Sarvam-1 scores higher on Hindi than Hinglish (+0.31 points, about 3 questions of 955) and OpenHathi higher on Romanized Hinglish (+0.73 points, about 7 questions) — small absolute differences on models that answer roughly one question in 16 to begin with. On general knowledge the pattern holds for all 16 models regardless.
Take both away and it falls off a cliff: only 41.4% (38.5–44.3) survives on general knowledge and 70.2% (68.4–71.8) on maths. Measured as a paired difference on the same questions — the right way to compare two forms, since two overlapping marginal intervals would not by themselves rule an effect out — the gap between Hinglish and Romanized Hindi is 45.1 points on general knowledge (95% CI 41.8 to 48.2) and 25.6 on maths (95% CI 23.8 to 27.4).
Which of the two changes costs more. Every question exists in all four combinations of where the words come from and which script they are written in, so the two effects can be separated directly:
| Hindi script | English letters | Cost of the script change | |
|---|---|---|---|
| Mixed with English words (Hinglish) | 86.5% | 70.3% | -16.2 |
| All Hindi | 74.2% | 41.4% | -32.8 |
| Cost of dropping the English words | -12.3 | -28.9 |
Read down the last column: switching to English letters costs 16.2 points in a mixed sentence but 32.8 in an all-Hindi one. The two factors compound rather than merely add: 16.6 points of extra loss on general knowledge (95% CI 13.5 to 19.8) and 7.9 on maths (95% CI 6.4 to 9.4). Both intervals exclude zero.
So the fair summary is not that mixing is harmless. It is that the script is the more expensive of the two changes, and that the English words in a Hinglish sentence do most of the work of keeping a model oriented once the script is gone.
↔ swipe to see the whole chart
Every number behind this chart is in the results tables.
This is the result we did not expect. Nemotron 3.5 Lightning scored the highest of all 16 models on English general knowledge — 89.36%. But when the same questions were typed in Hindi-in-English-letters, it kept only 47.3% of that ability, placing it 5th out of 15 — effectively tied with Llama 4 Scout, a difference of well under one question in 1,024.
Sarvam-M 24B scored 6.35 points lower in English — and kept 71.7%. Raw capability in English simply does not predict robustness here. What is associated with holding up is having been trained on Indian-language text.
| Model | English score | Floor-adj. retention |
|---|---|---|
| Nemotron 3.5 Lightning best English | 89.36% | 47.3% |
| Llama 4 Scout | 84.67% | 47.3% |
| Sarvam-M 24B | 83.01% | 71.7% |
| gpt-oss-20b | 82.13% | 61.9% |
| Sarvam-30B† | 82.13% | 69.6% |
| Phi-4 14B | 81.64% | 23.8% |
| Gemma 3 27B | 80.27% | 52.3% |
| Mistral Small 3.1 24B | 80.27% | 35.2% |
↔ swipe to see the whole chart
Every number behind this chart is in the results tables.
The most informative comparison is between two closely related checkpoints. Sarvam-M 24B is built on Mistral Small 3.1 24B — same parameter count, same base architecture — with additional Indian-language training.
| Test, Hindi-in-Latin | Mistral Small 24B | Sarvam-M 24B | Difference |
|---|---|---|---|
| General knowledge | 44.43% | 66.60% | +22.17 pts |
| Maths word problems | 69.95% | 83.98% | +14.03 pts |
Floor-adjusted retention rises from 35.2% to 71.7% — roughly double — at the same model size.
This comparison is consistent with a substantial benefit from Indic post-training, but it does not isolate which part of the recipe caused it. These are released checkpoints, not a controlled ablation: post-training data, instruction tuning, tokenizer and other implementation choices all differ alongside the Indic data. A causal attribution would need an ablation by the model’s authors.
↔ swipe to see the whole chart
Every number behind this chart is in the results tables.
Given romanized Hindi input, some models never commit to an answer — they keep writing until cut off. With no answer to extract, the question is scored wrong; it is never dropped from the total, so this failure is already paid for in the accuracy figures above. On maths problems it happens to 4.4% of questions in Hindi-in-Latin against 2.0% in English.
↔ swipe to see the whole chart
Every number behind this chart is in the results tables.
A bigger budget does not fix it — measured, not assumed. Whenever a run hit its token ceiling we re-ran just the affected rows with a larger budget, usually double, across 13 models; every attempt is still on disk. Of 3,047 maths rows that were cut off and then retried with more room, 1,890 (62%) still did not finish. The share that stayed unfinished is broadly similar in every language form (53–70%), so this is not a ceiling set slightly too low. These models do not stop.
Nor is it simply that romanized text is longer. Romanizing changes spelling, not word count, so a romanized question costs almost exactly what the same question costs in Hindi script:
| Hindi script | English letters | |
|---|---|---|
| Mixed with English words (Hinglish) | 1.12× | 1.11× |
| All Hindi | 1.26× | 1.23× |
Whatever romanized input is doing to these models, it is not merely making the input longer.
The main measure in this report is floor-adjusted retention — the share of a model’s own ability that survives when the question changes language.
It is not simply “score in Hindi ÷ score in English”, because that flatters weak models. In a four-option multiple-choice test, a model that knows nothing still scores about 25% by guessing. That 25% is the floor, and subtracting it first is what “floor-adjusted” means:
floor-adjusted retention = (score − floor) ÷ (English score − floor)
A real example. OpenHathi 7B scored 32.13% in English and 25.00% on Hindi-in-Latin-letters.
| Naive measure | 25.00 ÷ 32.13 | 77.8% | looks like it held up fine |
| Ability in English | 32.13 − 25 | 7.13 pts | above guessing |
| Ability in Hindi | 25.00 − 25 | 0.00 pts | above guessing |
| Floor-adjusted | 0.00 ÷ 7.13 | 0.0% | the truth: nothing survived |
The model scored exactly what guessing scores. So: 100% means the language change cost nothing, 44.4% (Gemma 3 12B on form 5) means it kept about half of what it knew, and 0% means it is down to guessing.
| Task | Floor used | Why |
|---|---|---|
| General knowledge | 25 | Four options per question, so guessing scores ~25%. |
| Maths word problems | 0 | The answer is a number the model has to work out. You cannot guess “7,412”, so there is no free score to subtract. |
With a floor of 0 the formula collapses to plain score ÷ English score. That is why the maths retention figures are simply the ratio of the two scores, while the general-knowledge ones are always lower than that ratio would suggest.
The figures in section 1 are unweighted means across eligible models of each model’s own floor-adjusted retention — per-model retention first, then averaged, so every model counts equally regardless of size.
Eligibility. A model is included for a task if its English score clears 35% on general knowledge (guess floor 25) or 15% on maths (no floor). That leaves 15 of 16 models on general knowledge and 13 of 16 on maths — the counts differ because more models sit near the floor on maths. Values are not clipped at 0 or 100; no eligible model produces a negative figure.
The cut-off is not load-bearing. Moving it shifts reported Romanized Hindi retention by at most 5.8 points across every threshold we tried, and never reorders the five forms. The full sweep is in the appendix.
Uncertainty. 95% intervals come from a paired bootstrap over questions (2,000 resamples, seed 1234). One resample of question indices is applied to every condition and model at once, so differences between conditions stay interpretable. These intervals quantify uncertainty from having sampled a finite set of questions, with the tested model set held fixed; they say nothing about how the result would generalise to other models.
| Task | Hinglish | Hindi | Romanized Hinglish | Romanized Hindi |
|---|---|---|---|---|
| General knowledge | 86.5% [84.2, 88.8] | 74.2% [71.5, 76.9] | 70.3% [67.6, 73.1] | 41.4% [38.5, 44.3] |
| Maths word problems | 95.7% [94.9, 96.6] | 88.1% [86.7, 89.5] | 85.7% [84.3, 87.0] | 70.2% [68.4, 71.8] |
Percentage of questions answered correctly. Bold marks each model’s English score — its own ceiling. The last column is how much of that survives the hardest form. Dimmed rows were already close to guessing in English, so they had nothing to lose.
† Sarvam-30B ran with its reasoning capped at 4,000 tokens — uncapped it left a quarter of rows unfinished and could not be scored — so its figures are accuracy under a reasoning cap, not with unconstrained reasoning.
| Model | English | Hinglish | Hindi | Romanized Hinglish | Romanized Hindi | Romanized Hindi loss vs English | Floor-adj. retention |
|---|---|---|---|---|---|---|---|
| Sarvam-M 24B | 83.01 | 79.69 | 74.22 | 74.02 | 66.60 | -16.41 | 71.7% |
| Sarvam-30B† | 82.13 | 78.12 | 75.39 | 73.14 | 64.75 | -17.38 | 69.6% |
| gpt-oss-20b | 82.13 | 77.73 | 75.68 | 74.41 | 60.35 | -21.78 | 61.9% |
| Gemma 3 27B | 80.27 | 72.36 | 69.34 | 66.02 | 53.91 | -26.37 | 52.3% |
| Nemotron 3.5 Lightning | 89.36 | 85.74 | 81.93 | 78.12 | 55.47 | -33.89 | 47.3% |
| Llama 4 Scout | 84.67 | 77.34 | 75.49 | 69.34 | 53.22 | -31.45 | 47.3% |
| Gemma 3 12B | 75.20 | 67.87 | 65.04 | 61.43 | 47.27 | -27.93 | 44.4% |
| Krutrim-2 12B | 64.36 | 56.15 | 52.73 | 50.88 | 42.19 | -22.17 | 43.7% |
| Gemma 3 12B (compressed) | 75.00 | 65.62 | 63.18 | 59.38 | 42.97 | -32.03 | 35.9% |
| Mistral Small 3.1 24B | 80.27 | 70.51 | 57.52 | 64.16 | 44.43 | -35.84 | 35.2% |
| Sarvam-1 2B | 49.80 | 47.07 | 43.95 | 38.87 | 32.52 | -17.29 | 30.3% |
| Llama 3.1 8B | 68.55 | 59.28 | 46.19 | 51.86 | 36.23 | -32.32 | 25.8% |
| Phi-4 14B | 81.64 | 74.02 | 64.26 | 59.57 | 38.48 | -43.16 | 23.8% |
| Gemma 3 4B | 63.67 | 57.03 | 50.68 | 48.44 | 32.03 | -31.64 | 18.2% |
| Param-1 2.9B | 47.56 | 44.14 | 36.72 | 36.62 | 28.22 | -19.34 | 14.3% |
| OpenHathi 7B | 32.13 | 29.10 | 23.14 | 26.86 | 25.00 | -7.13 | — |
↔ this table scrolls sideways — the model column stays fixed. Eight columns; the last is floor-adjusted retention.
| Model | English | Hinglish | Hindi | Romanized Hinglish | Romanized Hindi | Romanized Hindi loss vs English | Floor-adj. retention |
|---|---|---|---|---|---|---|---|
| Sarvam-M 24B | 96.34 | 94.35 | 89.42 | 91.52 | 83.98 | -12.36 | 87.2% |
| Gemma 3 27B | 95.92 | 94.14 | 88.80 | 89.21 | 81.57 | -14.35 | 85.0% |
| Llama 4 Scout | 96.44 | 94.76 | 89.53 | 88.69 | 79.06 | -17.38 | 82.0% |
| gpt-oss-20b | 96.34 | 93.72 | 89.21 | 88.59 | 78.95 | -17.38 | 82.0% |
| Gemma 3 12B | 94.35 | 91.83 | 85.97 | 85.65 | 75.81 | -18.53 | 80.4% |
| Gemma 3 12B (compressed) | 94.76 | 91.73 | 85.34 | 83.87 | 73.72 | -21.05 | 77.8% |
| Sarvam-30B† | 96.44 | 93.51 | 86.91 | 86.07 | 74.66 | -21.78 | 77.4% |
| Mistral Small 3.1 24B | 95.71 | 91.20 | 79.79 | 83.77 | 69.95 | -25.76 | 73.1% |
| Nemotron 3.5 Lightning | 96.75 | 95.18 | 89.42 | 85.24 | 63.98 | -32.77 | 66.1% |
| Krutrim-2 12B | 82.09 | 74.66 | 64.40 | 66.28 | 52.46 | -29.63 | 63.9% |
| Phi-4 14B | 96.13 | 92.77 | 85.13 | 77.59 | 48.80 | -47.33 | 50.8% |
| Gemma 3 4B | 90.26 | 86.28 | 77.70 | 66.18 | 40.63 | -49.63 | 45.0% |
| OpenHathi 7B | 6.28 | 5.13 | 3.46 | 5.86 | 2.83 | -3.46 | — |
| Llama 3.1 8B | 88.27 | 75.08 | 65.55 | 55.81 | 36.75 | -51.52 | 41.6% |
| Param-1 2.9B | 6.18 | 9.21 | 5.65 | 4.71 | 2.41 | -3.77 | — |
| Sarvam-1 2B | 8.59 | 5.65 | 5.97 | 3.25 | 2.93 | -5.65 | — |
↔ this table scrolls sideways — the model column stays fixed. Eight columns; the last is floor-adjusted retention.
Romanized Hindi loss vs English is how many percentage points the model dropped, measured against its own English score. Floor-adjusted retention is that same drop expressed as a share of the model’s real ability — see “how to read the numbers” above.
Choosing a model for Indian users? English benchmark scores alone can mislead you. The best English model here was mid-table once questions were typed in romanized form. Test on romanized input before committing.
Building a product? The input format matters as much as the model. Nudging users toward Hindi script, or converting romanized input before it reaches the model, may recover a large part of the loss — but note the ceiling: Hindi in its own script still retains only 74.2% on general knowledge and 88.1% on maths. Transliterating perfectly buys you the Hindi-script row, not the English one, and real transliteration carries its own accuracy, latency and maintenance costs on top.
Training a model? The gap can be materially reduced. Targeted training on Indian-language data is associated with roughly double the retention at the same model size — on a comparison of two released checkpoints rather than a controlled ablation.
The questions. Two standard test sets: a general-knowledge multiple-choice exam covering 57 subjects, and grade-school maths word problems. Every model saw identical questions in all five forms, so scores are directly comparable within a model.
Where each form came from. Only the Hinglish form is published research data. The other four were assembled for this study and aligned back to it question-by-question, so all five ask the same thing.
| Form | General knowledge | Maths word problems |
|---|---|---|
| Hinglish | CodeMixBench | CodeMixBench |
| English | cais/mmlu, rejoined by question id | openai/gsm8k, matched by its worked solution |
| Hindi | CohereLabs/Global-MMLU (Hindi) | bingbangboom/gsm8k-hindi (MIT) |
| Romanized Hinglish | Hinglish, transliterated | Hinglish, transliterated |
| Romanized Hindi | Hindi, transliterated | Hindi, transliterated |
Transliteration converts the script while leaving the words unchanged; English words inside a Hinglish sentence are left alone. Correct answers always come from the English originals, never from the translated datasets, so a mistranslated question cannot accidentally score as right — it simply fails.
Scoring. An answer is correct if it matches the known answer exactly. Models ran with settings that make them deterministic, so the same question gives the same answer every time. Where a model reasons before answering, only its final answer is scored; if it never reaches one before hitting its length limit, there is no answer to score and the question is marked wrong. Unfinished replies are never dropped from the total, so all five forms are scored over exactly the same questions — which is what makes comparing them like for like.
Fairness. Each model is compared only against itself, so a weaker model is not penalised for being weak and a stronger one gets no free credit. The measure is how much each loses. Full reproduction steps are in the repository README.
Correct answers always come from the English originals, never from the translated datasets — the Hindi answer fields are unreliable. But that means the Hindi conditions measure “can the model answer the English question as rendered in Hindi”, which is not quite “can the model do Hindi”. If a translation drifts, a model answering the Hindi question correctly is still marked wrong.
Coverage. Only 21.6% of the Hindi general-knowledge items (221 of 1024) are marked human-verified by Global-MMLU. The Hindi maths set is row-aligned, and its own answer field is malformed for ~27% of rows, which is why we do not use it.
Exclusions already applied. 61 of 1,016 maths items (6.0%) were dropped because the Hindi question did not carry the same numbers as the English one.
Measured effect. Comparing the same 15 models on the human-verified subset against the machine-translated one: +0.52 points on Hindi and -0.31 on Romanized Hindi — small, and absent on the form carrying the headline result. Questions that every eligible model gets right in English and wrong in Hindi number 1 of 1,024 (0.1%) and 8 of 955 (0.84%); the reverse direction is 0 in both.
Unquantified: semantic drift preserving the numbers, such as “gave away” becoming “received”. One confirmed case: a maths item where “1/4 as big as” became “1/4 bigger than” — all eligible models agreed on the same wrong answer, each having solved the mistranslated question correctly.
Romanizing is a deterministic transliteration of text we already have, so any translation error is identical on both sides and cancels exactly. On that comparison alone:
| Task | Romanizing Hinglish | Romanizing Hindi | Ratio |
|---|---|---|---|
| General knowledge | 7.10 pts | 15.58 pts | 2.2× |
| Maths word problems | 9.29 pts | 16.68 pts | 1.8× |
Losing the script costs roughly twice as much when there is no English to fall back on.
Romanized Hindi has no standard spelling: the same word is written several ways by different people, and often by the same person. Our two Latin-script conditions come from a single deterministic transform — indic_transliteration 2.3.82 (DEVANAGARI -> ITRANS), then final-schwa deletion, anusvara/candrabindu -> n, visarga -> h, danda -> full stop and lowercasing (scripts/build_conditions.py:to_hinglish). That makes the conditions exactly reproducible, and it also makes them one point in a wide space of things people really type.
The transform has a known defect, and it is common. It does not model internal schwa deletion, so it writes men where a person writes mein, and kitane for kitne. At least one such form appears in 66–93% of rows across the four romanized sets:
| Romanized condition | Rows with a known-wrong form | Share of words |
|---|---|---|
mmlu_hineng_rom | 679 of 1,024 (66.3%) | 2.70% |
mmlu_hineng_hirom | 816 of 1,024 (79.7%) | 3.99% |
gsm8k_hineng_rom | 680 of 955 (71.2%) | 3.38% |
gsm8k_hineng_hirom | 888 of 955 (93.0%) | 6.13% |
Which way this biases the result. Our romanized text is slightly less natural than real typing, so it is plausibly further out of a model’s distribution than what a user would actually send. The romanization penalty reported here is therefore best read as an upper bound on the penalty for well-formed romanized Hindi. Pushing the other way, real input is more variable than ours — inconsistent spelling inside a single message — which a uniform scheme does not test at all.
Not validated against human-typed text. We did not collect human romanizations of these questions, so we cannot report agreement with them and have not claimed to. A seeded 64-row sample of the transform’s output beside its Devanagari source ships as results/romanization_sample.csv for anyone who reads Hindi to check by eye. Comparing against human-typed romanized Hindi is the clearest next step, and would tighten that bound.
A model scoring near the guessing floor in English cannot rank language forms, so models below a threshold are left out of the headline means. Any threshold invites the question of whether it was picked to flatter the result, so here is the whole sweep:
| Cut-off | Eligible | English | Hinglish | Hindi | Romanized Hinglish | Romanized Hindi |
|---|---|---|---|---|---|---|
| General knowledge, ≥25% | 16 | 100.0% | 84.7% | 68.0% | 67.6% | 38.9% |
| General knowledge, ≥30% | 16 | 100.0% | 84.7% | 68.0% | 67.6% | 38.9% |
| General knowledge, ≥35% used | 15 | 100.0% | 86.5% | 74.2% | 70.3% | 41.4% |
| General knowledge, ≥40% | 15 | 100.0% | 86.5% | 74.2% | 70.3% | 41.4% |
| General knowledge, ≥45% | 15 | 100.0% | 86.5% | 74.2% | 70.3% | 41.4% |
| General knowledge, ≥50% | 13 | 100.0% | 86.4% | 75.8% | 72.9% | 44.4% |
| Maths word problems, ≥5% | 16 | 100.0% | 96.3% | 85.1% | 82.6% | 64.4% |
| Maths word problems, ≥10% | 13 | 100.0% | 95.7% | 88.1% | 85.7% | 70.2% |
| Maths word problems, ≥15% used | 13 | 100.0% | 95.7% | 88.1% | 85.7% | 70.2% |
| Maths word problems, ≥20% | 13 | 100.0% | 95.7% | 88.1% | 85.7% | 70.2% |
| Maths word problems, ≥30% | 13 | 100.0% | 95.7% | 88.1% | 85.7% | 70.2% |
| Maths word problems, ≥40% | 13 | 100.0% | 95.7% | 88.1% | 85.7% | 70.2% |
| Decoding | greedy (temperature 0), fixed seed — deterministic |
| Precision | bfloat16 for all models except the two below |
| Quantized models | Gemma 3 12B (compressed) is 4-bit; Llama 4 Scout is a 4-bit weight-quantized release |
| Token budgets | per model, set by its context window and whether it reasons before answering; recorded per run |
| Answer parsing | shape-aware: handles models that answer first and models that reason first. A reply containing no answer — cut off mid-reasoning, say — is scored wrong, never dropped, so every form is scored over the same questions |
| Reasoning models | Sarvam-30B ran with reasoning capped at 4,000 tokens; uncapped it left a quarter of rows unfinished |
| Inference | vLLM, offline batch |
Exact checkpoint ids, per-model configuration, prompts, token budgets and the answer parser are all in the accompanying code repository; each run also writes a run_meta.json recording the settings actually used.
Hinglish questions from CodeMixBench (Yang & Chai, EMNLP 2025). English and Hindi matched question-by-question from cais/mmlu, openai/gsm8k, CohereLabs/Global-MMLU and bingbangboom/gsm8k-hindi. The alignment across all five forms, the two Latin-letter forms, the retention measure and the 16-model comparison are this study’s.