Writing India report
Every model is asked to paint the same kind of scene: an everyday Indian setting with a specific sentence rendered onto a signboard, in a specific Indic script. Nine open-weight models, the identical 500 prompts, the same three scorers throughout. Eight of the nine produce a flat 0.0% exact match on Hindi, Bengali, Marathi, Telugu and Kannada signage — regardless of parameter count or release date. The one exception, Z-Image-Turbo, is substantial only on Devanagari (Hindi and Marathi), marginal on Bengali (10%), and a flat 0% on both Dravidian scripts it was tested on.
This benchmark asks one concrete question: if a model is asked to paint a specific sentence onto a signboard inside an everyday Indian scene, in a named Indic script, does the sentence come out legible — and does the scene around it actually look Indian, rather than a generic backdrop with Indian props scattered on top? Two independent failures are possible, and the four scoring tracks below (six reported metrics, grouped into four tracks — see metrics) exist because collapsing them into one number would hide which one a given model is making: a model can render an authentic street while leaving its sign as gibberish, or it can spell the sign correctly onto a scene that reads as a Western stock photo with a dupatta added.
500 LLM-generated prompts across six categories and seven Indian macro-regions. 300 of them ask for a signboard bearing a specific string, split evenly — 50 apiece — across Hindi, Marathi, Bengali, Telugu, Kannada and English.
Every prompt also carries an LLM-generated expected_attributes
field: a checklist of 4–6 concrete, checkable details the scene should contain, phrased as a specific
claim rather than a general impression — "weighing scale at the shop counter," not "looks Indian." This
is the field Attribute Similarity and Attribute Recall are scored against, one cosine comparison per listed
attribute, independently of whatever score the full prompt sentence gets — so a model can be graded on
individual details, not just overall gist.
No existing corpus pairs "generate this specific string, in this specific Indic script, on a photorealistic Indian scene" with a ground-truth answer — that pairing only exists if someone writes it. Authoring also buys a control that no photograph collection can: the 300 Indic Text prompts are built as a parallel core, the same handful of scenes (a bookshop, a vegetable stall, a ticket counter, etc.) requested with the identical sentence structure in all six languages. That holds scene difficulty constant across languages, so a score gap between Hindi and Kannada can be attributed to script rendering rather than to one scene happening to be harder to draw. Six real photographs of that same shopfront, in six different scripts, do not exist to be collected.
A handful of adjacent benchmarks exist — some test whether a model understands a prompt written in an Indic language, others test reading comprehension over real photographs that already contain text. None score whether a model can render a specific Indic script legibly inside an otherwise-authentic generated scene; that gap is what these 500 prompts were written to fill.
| by category | |
|---|---|
| Indic Text (a sign to render) | 300 |
| People & Professions | 62 |
| Domestic Environment | 38 |
| Public Environment | 38 |
| Food & Objects | 38 |
| Clothing & Appearance | 24 |
| by region | |
|---|---|
| South India | 142 |
| Pan-India | 98 |
| North India | 80 |
| West / East India | 75 each |
| Central India | 18 |
| Northeast India | 12 |
Non-text prompts still carry a 4–6 item attribute checklist (here: a farmer, a wheat field, mid-day light, plausible tools and dress) — Prompt Alignment and Attribute Recall apply to all 500 prompts; only the 300 Indic Text prompts carry an OCR/CER track.
Six metrics below, grouped into four evaluation tracks: Prompt Fidelity (prompt alignment + attribute similarity + attribute recall, all from SigLIP 2), Cultural Authenticity (Qwen3-VL-32B), OCR Exact Match and CER (both IndicPhotoOCR). The first track is split into three numbers because a single cosine score cannot separate "the whole scene matches" from "this one attribute is present" — see the worked examples for why that distinction matters.
| model | prompt align | attr sim | attr recall | cultural auth | ocr exact match | cer |
|---|---|---|---|---|---|---|
| FLUX.2 [dev] | 0.204 | 0.096 | 93.2% | 4.60 | 15.3% | 0.541 |
| Z-Image-Turbo | 0.195 | 0.093 | 91.3% | 4.39 | 35.7% | 0.342 |
| SD 3.5 Large | 0.198 | 0.095 | 91.6% | 4.26 | 11.0% | 0.645 |
| FLUX.2 [klein] 4B | 0.180 | 0.092 | 92.5% | 4.15 | 8.7% | 0.658 |
| Qwen-Image-2512 | 0.219 | 0.094 | 89.7% | 3.97 | 16.0% | 0.651 |
| PixArt-Σ | 0.238 | 0.089 | 84.1% | 3.46 | 0.0% | 0.882 |
| Lumina-Image 2.0 | 0.242 | 0.095 | 86.9% | 3.25 | 1.7% | 0.744 |
| SANA 1.5 4.8B | 0.238 | 0.090 | 84.6% | 3.25 | 1.7% | 0.785 |
| HiDream-I1-Dev | 0.201 | 0.077 | 80.8% | 2.62 | 9.7% | 0.682 |
Sorted by Cultural Authenticity. Bold marks the best figure per column. FLUX.2 [dev] leads general quality; Z-Image-Turbo is a full class apart on the two text-specific columns and a full class behind on none of the others.
| model | devanagari (hi) | devanagari (mr) | bengali | telugu | kannada | latin (en) | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EM | CER | EM | CER | EM | CER | EM | CER | EM | CER | EM | CER | |
| Z-Image-Turbo | 68.0% | 0.055 | 54.0% | 0.065 | 10.0% | 0.290 | 0.0% | 0.817 | 0.0% | 0.808 | 82.0% | 0.030 |
| FLUX.2 [dev] | 0.0% | 0.474 | 0.0% | 0.497 | 0.0% | 0.684 | 0.0% | 0.811 | 0.0% | 0.801 | 92.0% | 0.015 |
| Qwen-Image-2512 | 0.0% | 0.702 | 0.0% | 0.754 | 0.0% | 0.781 | 0.0% | 0.876 | 0.0% | 0.843 | 96.0% | 0.005 |
| SD 3.5 Large | 0.0% | 0.763 | 0.0% | 0.780 | 0.0% | 0.743 | 0.0% | 0.784 | 0.0% | 0.808 | 66.0% | 0.050 |
| HiDream-I1-Dev | 0.0% | 0.772 | 0.0% | 0.807 | 0.0% | 0.786 | 0.0% | 0.832 | 0.0% | 0.835 | 58.0% | 0.113 |
| FLUX.2 [klein] 4B | 0.0% | 0.775 | 0.0% | 0.757 | 0.0% | 0.750 | 0.0% | 0.815 | 0.0% | 0.787 | 52.0% | 0.113 |
| Lumina-Image 2.0 | 0.0% | 0.820 | 0.0% | 0.789 | 0.0% | 0.817 | 0.0% | 0.796 | 0.0% | 0.900 | 10.0% | 0.380 |
| SANA 1.5 4.8B | 0.0% | 0.848 | 0.0% | 0.808 | 0.0% | 0.831 | 0.0% | 0.838 | 0.0% | 0.896 | 10.0% | 0.517 |
| PixArt-Σ | 0.0% | 0.885 | 0.0% | 0.901 | 0.0% | 0.916 | 0.0% | 0.861 | 0.0% | 0.964 | 0.0% | 0.782 |
The dot marks a flat 0.0% exact match — 42 of the 45 Indic-script EM cells in this table (43 of all 54, including English). Every cell is 50 prompts; at that sample size a true 0% cannot be distinguished from a small non-zero rate — the 95% Wilson interval on an observed 0/50 is [0.0%, 7.1%], not a point estimate of exactly zero. It does distinguish cleanly from Z-Image-Turbo's non-zero cells: its weakest, Bengali at 10.0%, has a 95% CI of [4.3%, 21.4%], which does not overlap 0%. But the CER column beside EM is never uniform even where EM is: FLUX.2 [dev] sits closer to legible on Hindi/Marathi (CER ≈0.48–0.50) than PixArt-Σ does anywhere (CER 0.86–0.96). Exact match cannot tell these apart; CER is what shows a near-miss is not the same failure as noise.
Neither automated scorer was taken on trust. Both were cross-checked against human review on a stratified sample of 100 prompts, scored for all nine models — 900 images, rated blind to the automated verdicts. Cultural Authenticity was re-scored by a reviewer on the same 0–5 scale the judge uses. Text rendering was reviewed by a language expert fluent in Telugu and English, who rated how legible the script on each sign actually was; those are the two languages that expert can read, which is what the legibility check covers.
| model | cultural authenticity (n = 100) | english legibility (n = 25) | telugu legibility (n = 25) | |||
|---|---|---|---|---|---|---|
| expert | judge | expert | ocr em | expert | ocr em | |
| Qwen-Image-2512 | 4.81 | 4.13 | 5.00 | 96% | 0.00 | 0% |
| FLUX.2 [dev] | 4.74 | 4.63 | 5.00 | 84% | 0.00 | 0% |
| SD 3.5 Large | 4.50 | 4.47 | 4.68 | 72% | 0.00 | 0% |
| FLUX.2 [klein] 4B | 4.45 | 4.27 | 4.20 | 52% | 0.00 | 0% |
| Z-Image-Turbo | 4.41 | 4.30 | 4.84 | 96% | 0.00 | 0% |
| Lumina-Image 2.0 | 3.67 | 3.14 | 3.40 | 20% | 0.00 | 0% |
| SANA 1.5 4.8B | 3.61 | 3.24 | 2.56 | 20% | 0.00 | 0% |
| PixArt-Σ | 3.28 | 3.42 | 0.00 | 0% | 0.00 | 0% |
| HiDream-I1-Dev | 2.93 | 2.66 | 4.64 | 60% | 0.00 | 0% |
Judge scores are recomputed over the same 100 sampled prompts, so the columns compare like with like. The judge holds up directionally, not precisely: it matches the human exactly on 43.0% of images, correlates at r = 0.36, and reads slightly stricter on eight of nine models — the widest gap being Qwen-Image-2512 (human 4.81, judge 4.13). The ordering survives — the same models lead, HiDream-I1-Dev trails either way — so Cultural Authenticity is a dependable ranking rather than a calibrated absolute score. The legibility check is much cleaner: the expert's English ratings move almost in lockstep with OCR exact match, from 5.00 / 96% at the top to a matching 0.00 / 0% for PixArt-Σ at the bottom. And the Telugu column is a flat zero for every model, Z-Image-Turbo included — a native reader confirming, independently of any OCR system, that not one of the nine produced legible Telugu.
Prompt:
Expected text:
पुस्तक भंडार — "Pustak Bhandar", Book Depot
None of the four is an exact match — the metric behind the matrix above is binary. But CER shows the texture behind that binary: Z-Image-Turbo and FLUX.2 [dev] both land one glyph away from correct; Qwen-Image-2512 gets half the glyphs right; SD 3.5 Large's sign is Devanagari-shaped but unreadable end to end. A pass/fail metric alone would score all four the same 0%; CER is why both figures are reported.
Prompt:
Expected attributes: dairy farmer at work · milking shed with cattle · milk collection vessels · working farm setting · western Indian dairy farming context.
All four hit every expected attribute here — cattle, milk vessels, a working shed — and three of four score a perfect 5/5 authenticity. Qwen-Image-2512 is marked down slightly to 4/5: the judge's rationale is specific rather than a vibe, flagging the milking pail as "a modern metal cup with a handle… more typical of Western or urban dairy setups rather than traditional Indian hand-milking practices." One anachronistic object, correctly named, cost exactly one point.
Prompt:
“A small vegetable shop in northern India with authentic details: a weathered blue painted signboard with Hindi script, a corrugated metal awning, and fresh vegetables displayed on metal trays … all consistent with local market practices.”
This is what the other examples are measured against: Z-Image-Turbo on a shorter, more common phrase (three words, "Vegetable Shop") — every track lands well, all at once, on the same image. Across all nine models and the full 300-prompt text-rendering set (English included), only 299 image/model pairs land an exact OCR match out of 2,700 attempted — the bookshop example above, not this one, is the far more common outcome.
Z-Image-Turbo is the exception among the nine models scored here: 68.0% exact match on Hindi, 54.0% on Marathi (both Devanagari), a much weaker 10.0% on Bengali, and a flat 0.0% on both Telugu and Kannada — the two Dravidian scripts tested. Its capability is specific to one script family, not a general Indic solution, and even inside that family it is partial.
Qwen-Image-2512 renders English at 96.0% and FLUX.2 [dev] at 92.0% — the two best in the field. Both score exactly 0.0% on every Indic script, identical to Lumina-Image 2.0 and SANA 1.5, whose English is a mere 10.0%. Eight of nine models land at a flat 0% on Devanagari, Bengali, Telugu and Kannada regardless of where their English score falls, which is what a separately trained capability looks like rather than one general "text-rendering quality" dial.
On the shared example earlier in this report, FLUX.2 [dev] rendered the requested "पुस्तक भंडार" as "पुत्तक भंडार" — one wrong glyph out of twelve (CER 0.083), the kind of slip a reader would likely still parse. SD 3.5 Large's reading of the same sign (CER 0.833) is not Devanagari-shaped text a reader could parse at all. Exact match alone scores both a flat 0% and cannot tell them apart; CER shows the first is a model that can render the script and mostly did, just imperfectly, while the second genuinely cannot.
This benchmark scores outputs, not training corpora, so this is an explanation rather than a measurement. Web-scraped caption corpora are dominated by English captions and Latin-script signage, and the models that render Chinese or Japanese reliably got there through corpora deliberately assembled for it. No comparable public corpus pairs images with legible Devanagari, Bengali, Telugu or Kannada text — which would mean a model fails these scripts not from weaker text-rendering skill, but from rarely having seen one.
Wording, instruction language and examples can all move a generative model's output. This benchmark fixes one sentence template across all 500 prompts and all nine models specifically so a score gap reflects a difference between models rather than a difference in how each one happened to be asked — the same prompt set can be reused as-is to test a tenth model against the exact same bar.
run_metadata.json, written before generation starts