9 MODELS · 500 PROMPTS · 5 INDIC LANGUAGES / 4 SCRIPTS TESTED
~/india-in-the-frame

Writing India report

9 open-weight T2I models500 prompts4 Indic scripts

Every model is asked to paint the same kind of scene: an everyday Indian setting with a specific sentence rendered onto a signboard, in a specific Indic script. Nine open-weight models, the identical 500 prompts, the same three scorers throughout. Eight of the nine produce a flat 0.0% exact match on Hindi, Bengali, Marathi, Telugu and Kannada signage — regardless of parameter count or release date. The one exception, Z-Image-Turbo, is substantial only on Devanagari (Hindi and Marathi), marginal on Bengali (10%), and a flat 0% on both Dravidian scripts it was tested on.

report · SigLIP2 + Qwen3-VL-32B + IndicPhotoOCR · 9 models, 6 languages · Sept 2026
~/india-in-the-frame

Writing India

Nine open-weight text-to-image models, scored on the identical 500-prompt set for cultural authenticity, prompt fidelity and legible Indic-script rendering · Sept 2026

the task

This benchmark asks one concrete question: if a model is asked to paint a specific sentence onto a signboard inside an everyday Indian scene, in a named Indic script, does the sentence come out legible — and does the scene around it actually look Indian, rather than a generic backdrop with Indian props scattered on top? Two independent failures are possible, and the four scoring tracks below (six reported metrics, grouped into four tracks — see metrics) exist because collapsing them into one number would hide which one a given model is making: a model can render an authentic street while leaving its sign as gibberish, or it can spell the sign correctly onto a scene that reads as a Western stock photo with a dupatta added.

the dataset

500 LLM-generated prompts across six categories and seven Indian macro-regions. 300 of them ask for a signboard bearing a specific string, split evenly — 50 apiece — across Hindi, Marathi, Bengali, Telugu, Kannada and English.

Every prompt also carries an LLM-generated expected_attributes field: a checklist of 4–6 concrete, checkable details the scene should contain, phrased as a specific claim rather than a general impression — "weighing scale at the shop counter," not "looks Indian." This is the field Attribute Similarity and Attribute Recall are scored against, one cosine comparison per listed attribute, independently of whatever score the full prompt sentence gets — so a model can be graded on individual details, not just overall gist.

No existing corpus pairs "generate this specific string, in this specific Indic script, on a photorealistic Indian scene" with a ground-truth answer — that pairing only exists if someone writes it. Authoring also buys a control that no photograph collection can: the 300 Indic Text prompts are built as a parallel core, the same handful of scenes (a bookshop, a vegetable stall, a ticket counter, etc.) requested with the identical sentence structure in all six languages. That holds scene difficulty constant across languages, so a score gap between Hindi and Kannada can be attributed to script rendering rather than to one scene happening to be harder to draw. Six real photographs of that same shopfront, in six different scripts, do not exist to be collected.

A handful of adjacent benchmarks exist — some test whether a model understands a prompt written in an Indic language, others test reading comprehension over real photographs that already contain text. None score whether a model can render a specific Indic script legibly inside an otherwise-authentic generated scene; that gap is what these 500 prompts were written to fill.

Prompt count by category, out of 500 total
by category
Indic Text (a sign to render)300
People & Professions62
Domestic Environment38
Public Environment38
Food & Objects38
Clothing & Appearance24
Prompt count by Indian region, out of 500 total
by region
South India142
Pan-India98
North India80
West / East India75 each
Central India18
Northeast India12

one record per category

{"id": "script_kn_032", "category": "Indic Text", "subcategory": "Restaurant Sign", "prompt": "Generate a photorealistic photograph of a breakfast stall in a town in southern India, with a painted board above the stall displaying the Kannada text: ಇಡ್ಲಿ ವಡೆ", "language": "Kannada", "script": "Kannada", "region": "South India", "expected_text": "ಇಡ್ಲಿ ವಡೆ", "expected_attributes": ["photorealistic southern Indian breakfast stall", "idli steamer and vada frying pan", "chutney containers on the counter", "Kannada script on the painted board", "clearly readable requested text"], "evaluation_tracks": ["Prompt Alignment", "OCR Exact Match", "CER", "Cultural Authenticity"], "text_visibility": "prominent", "notes": "Language-specific scenario: signage of this kind is characteristic of South India and is not part of the parallel core. The requested string should be the only text on the painted board so that OCR Exact Match and CER stay unambiguous. Cultural Authenticity is judged on the surrounding scene rather than on the sign itself."}
{"id": "prof_001", "category": "People & Professions", "subcategory": "Farmer", "prompt": "Generate a photorealistic image of a farmer working in a wheat field in North India during the day.", "language": null, "script": null, "region": "North India", "expected_text": null, "expected_attributes": ["farmer working in the field", "wheat crop at a plausible growth stage", "farming activity in progress", "daytime outdoor setting", "North Indian agricultural context"], "evaluation_tracks": ["Prompt Alignment", "Cultural Authenticity"], "text_visibility": null, "notes": "Judge the crop and the work being done. Mechanised and non-mechanised methods are both plausible."}
{"id": "dom_001", "category": "Domestic Environment", "subcategory": "Indian Kitchen", "prompt": "Generate a photorealistic photograph of the kitchen of an ordinary middle-class home in South India, showing the cookware, utensils, storage containers and ingredients that are in everyday use.", "language": null, "script": null, "region": "South India", "expected_text": null, "expected_attributes": ["middle-class household kitchen", "Indian cookware in everyday use", "Indian utensils", "storage containers for staples and spices", "plausible food ingredients", "South Indian household context"], "evaluation_tracks": ["Prompt Alignment", "Cultural Authenticity"], "text_visibility": null, "notes": "The failure mode is a Western kitchen with a few Indian objects added. Judge the working layout, the storage habits and whether South Indian specifics survive rather than collapsing into a generic pan-Indian look."}
{"id": "pub_001", "category": "Public Environment", "subcategory": "Street Market", "prompt": "Generate a photorealistic photograph of a neighbourhood vegetable and fruit market in a North Indian city in the morning, with vendors at their stalls and customers buying produce.", "language": null, "script": null, "region": "North India", "expected_text": null, "expected_attributes": ["neighbourhood produce market", "vegetable and fruit stalls", "vendors working at the stalls", "customers buying produce", "morning daylight", "North Indian urban street context"], "evaluation_tracks": ["Prompt Alignment", "Cultural Authenticity"], "text_visibility": null, "notes": "Judge how the market works: stall construction, how produce is heaped, weighing equipment and the produce mix. A working market is busy by nature, so activity is prompt-aligned."}
{"id": "food_001", "category": "Food & Objects", "subcategory": "South Indian Breakfast", "prompt": "Generate a photorealistic close-up photograph of a typical South Indian breakfast served at home on a stainless steel plate, with the usual accompaniments in small bowls beside it.", "language": null, "script": null, "region": "South India", "expected_text": null, "expected_attributes": ["recognisable South Indian breakfast items", "stainless steel plate", "small bowls for accompaniments", "plausible accompaniments", "home dining surface"], "evaluation_tracks": ["Prompt Alignment", "Cultural Authenticity"], "text_visibility": null, "notes": "Serving ware is part of the ground truth. The central failure is a South Indian breakfast rendered as generic North Indian curry or as a restaurant platter."}
{"id": "cloth_001", "category": "Clothing & Appearance", "subcategory": "Everyday Clothing", "prompt": "Generate a photorealistic image of a group of adults waiting at a city bus stop in India on an ordinary weekday morning, dressed as people typically dress for work and daily errands.", "language": null, "script": null, "region": "Pan-India", "expected_text": null, "expected_attributes": ["adults at a city bus stop", "everyday work and errand clothing", "plausible mix of Indian and Western everyday dress", "bags and items carried on a commute", "Indian urban street context"], "evaluation_tracks": ["Prompt Alignment", "Cultural Authenticity"], "text_visibility": null, "notes": "An ordinary weekday is requested, so festive dress is an unrequested addition. Both Indian and Western everyday clothing are accurate; accept a realistic mix rather than a single style."}

Non-text prompts still carry a 4–6 item attribute checklist (here: a farmer, a wheat field, mid-day light, plausible tools and dress) — Prompt Alignment and Attribute Recall apply to all 500 prompts; only the 300 Indic Text prompts carry an OCR/CER track.

metrics

Six metrics below, grouped into four evaluation tracks: Prompt Fidelity (prompt alignment + attribute similarity + attribute recall, all from SigLIP 2), Cultural Authenticity (Qwen3-VL-32B), OCR Exact Match and CER (both IndicPhotoOCR). The first track is split into three numbers because a single cosine score cannot separate "the whole scene matches" from "this one attribute is present" — see the worked examples for why that distinction matters.

prompt alignment[−1, 1]
SigLIP2 cosine similarity between the image's embedding and the full prompt sentence's embedding, both produced by the same contrastive vision-language encoder.
attribute similarity[−1, 1]
Cosine similarity computed once per expected attribute given in the dataset against the image and then averaged, rather than once on the whole prompt sentence.
attribute recall[0, 100]%
Fraction of expected attributes whose cosine score clears a fixed threshold (0.04), calibrated in advance to separate attributes that are genuinely present from deliberately mismatched ones.
cultural authenticityinteger, [1, 5]
A vision-language model (Qwen3-VL-32B-Instruct) scores one image at a time against a five-point rubric anchored on a specific failure mode. The score covers only the scene — any rendered signage is judged separately, by OCR Exact Match and CER.
ocr exact match[0, 100]%
Whether the normalised text extracted from the image is identical to the normalised ground-truth string, binary per image, then averaged.
Character Error Rate[0, ∞)
The Levenshtein edit distance between extracted and ground-truth text, summed across every scored image and divided by the summed length of every ground-truth string. Lower is better, and unlike exact match it distinguishes a near-miss reading from a total miss.

results

Benchmark results by model, across all six metrics, sorted by Cultural Authenticity
modelprompt alignattr simattr recallcultural authocr exact matchcer
FLUX.2 [dev]0.2040.09693.2%4.6015.3%0.541
Z-Image-Turbo0.1950.09391.3%4.3935.7%0.342
SD 3.5 Large0.1980.09591.6%4.2611.0%0.645
FLUX.2 [klein] 4B0.1800.09292.5%4.158.7%0.658
Qwen-Image-25120.2190.09489.7%3.9716.0%0.651
PixArt-Σ0.2380.08984.1%3.460.0%0.882
Lumina-Image 2.00.2420.09586.9%3.251.7%0.744
SANA 1.5 4.8B0.2380.09084.6%3.251.7%0.785
HiDream-I1-Dev0.2010.07780.8%2.629.7%0.682

Sorted by Cultural Authenticity. Bold marks the best figure per column. FLUX.2 [dev] leads general quality; Z-Image-Turbo is a full class apart on the two text-specific columns and a full class behind on none of the others.

ocr exact match & CER, by script

OCR Exact Match and Character Error Rate by model and script, including English
model devanagari (hi) devanagari (mr) bengali telugu kannada latin (en)
EMCER EMCER EMCER EMCER EMCER EMCER
Z-Image-Turbo68.0%0.05554.0%0.06510.0%0.2900.0%0.8170.0%0.80882.0%0.030
FLUX.2 [dev]0.0%0.4740.0%0.4970.0%0.6840.0%0.8110.0%0.80192.0%0.015
Qwen-Image-25120.0%0.7020.0%0.7540.0%0.7810.0%0.8760.0%0.84396.0%0.005
SD 3.5 Large0.0%0.7630.0%0.7800.0%0.7430.0%0.7840.0%0.80866.0%0.050
HiDream-I1-Dev0.0%0.7720.0%0.8070.0%0.7860.0%0.8320.0%0.83558.0%0.113
FLUX.2 [klein] 4B0.0%0.7750.0%0.7570.0%0.7500.0%0.8150.0%0.78752.0%0.113
Lumina-Image 2.00.0%0.8200.0%0.7890.0%0.8170.0%0.7960.0%0.90010.0%0.380
SANA 1.5 4.8B0.0%0.8480.0%0.8080.0%0.8310.0%0.8380.0%0.89610.0%0.517
PixArt-Σ0.0%0.8850.0%0.9010.0%0.9160.0%0.8610.0%0.9640.0%0.782
EM (higher better)0.0%10%30–54%55–70%80%+
CER (lower better)<0.100.10–0.300.30–0.550.55–0.750.75–0.850.85+

The dot marks a flat 0.0% exact match — 42 of the 45 Indic-script EM cells in this table (43 of all 54, including English). Every cell is 50 prompts; at that sample size a true 0% cannot be distinguished from a small non-zero rate — the 95% Wilson interval on an observed 0/50 is [0.0%, 7.1%], not a point estimate of exactly zero. It does distinguish cleanly from Z-Image-Turbo's non-zero cells: its weakest, Bengali at 10.0%, has a 95% CI of [4.3%, 21.4%], which does not overlap 0%. But the CER column beside EM is never uniform even where EM is: FLUX.2 [dev] sits closer to legible on Hindi/Marathi (CER ≈0.48–0.50) than PixArt-Σ does anywhere (CER 0.86–0.96). Exact match cannot tell these apart; CER is what shows a near-miss is not the same failure as noise.

human validation

Neither automated scorer was taken on trust. Both were cross-checked against human review on a stratified sample of 100 prompts, scored for all nine models — 900 images, rated blind to the automated verdicts. Cultural Authenticity was re-scored by a reviewer on the same 0–5 scale the judge uses. Text rendering was reviewed by a language expert fluent in Telugu and English, who rated how legible the script on each sign actually was; those are the two languages that expert can read, which is what the legibility check covers.

Human ratings against the automated scorers, per model, on the same 100-prompt sample
modelcultural authenticity (n = 100)english legibility (n = 25)telugu legibility (n = 25)
expertjudgeexpertocr emexpertocr em
Qwen-Image-25124.814.135.0096%0.000%
FLUX.2 [dev]4.744.635.0084%0.000%
SD 3.5 Large4.504.474.6872%0.000%
FLUX.2 [klein] 4B4.454.274.2052%0.000%
Z-Image-Turbo4.414.304.8496%0.000%
Lumina-Image 2.03.673.143.4020%0.000%
SANA 1.5 4.8B3.613.242.5620%0.000%
PixArt-Σ3.283.420.000%0.000%
HiDream-I1-Dev2.932.664.6460%0.000%

Judge scores are recomputed over the same 100 sampled prompts, so the columns compare like with like. The judge holds up directionally, not precisely: it matches the human exactly on 43.0% of images, correlates at r = 0.36, and reads slightly stricter on eight of nine models — the widest gap being Qwen-Image-2512 (human 4.81, judge 4.13). The ordering survives — the same models lead, HiDream-I1-Dev trails either way — so Cultural Authenticity is a dependable ranking rather than a calibrated absolute score. The legibility check is much cleaner: the expert's English ratings move almost in lockstep with OCR exact match, from 5.00 / 96% at the top to a matching 0.00 / 0% for PixArt-Σ at the bottom. And the Telugu column is a flat zero for every model, Z-Image-Turbo included — a native reader confirming, independently of any OCR system, that not one of the nine produced legible Telugu.

worked example

Text rendering, compared — four models, one requested sign

Prompt:

Generate a photorealistic photograph of a small bookshop in a town in northern India, with a painted signboard above the entrance displaying the Hindi text: पुस्तक भंडार

Expected text:

पुस्तक भंडार — "Pustak Bhandar", Book Depot

Z-Image-Turbo generated bookshop signboard
Z-Image-Turbo
पुस्तक भंदार
no exact match · CER 0.083
one glyph off (ड→द)
FLUX.2 dev generated bookshop signboard
FLUX.2 [dev]
पुत्तक भंडार
no exact match · CER 0.083
one glyph off (स→त)
Qwen-Image-2512 generated bookshop signboard
Qwen-Image-2512
पुंतिक अदार
no exact match · CER 0.500
half the glyphs wrong
Stable Diffusion 3.5 Large generated bookshop signboard
SD 3.5 Large
no legible match
no exact match · CER 0.833
Devanagari-shaped, unreadable

None of the four is an exact match — the metric behind the matrix above is binary. But CER shows the texture behind that binary: Z-Image-Turbo and FLUX.2 [dev] both land one glyph away from correct; Qwen-Image-2512 gets half the glyphs right; SD 3.5 Large's sign is Devanagari-shaped but unreadable end to end. A pass/fail metric alone would score all four the same 0%; CER is why both figures are reported.

Cultural authenticity, without any text — four models, no script to render

Prompt:

Generate a photorealistic image of a dairy farmer at work in a milking shed in West India.

Expected attributes: dairy farmer at work · milking shed with cattle · milk collection vessels · working farm setting · western Indian dairy farming context.

Z-Image-Turbo generated dairy farmer in a milking shed
Z-Image-Turbo
5/5 attributes · auth 5/5
prompt align 0.130
FLUX.2 dev generated dairy farmer in a milking shed
FLUX.2 [dev]
5/5 attributes · auth 5/5
prompt align 0.121
Qwen-Image-2512 generated dairy farmer in a milking shed
Qwen-Image-2512
5/5 attributes · auth 4/5
prompt align 0.222
Stable Diffusion 3.5 Large generated dairy farmer in a milking shed
SD 3.5 Large
5/5 attributes · auth 5/5
prompt align 0.168

All four hit every expected attribute here — cattle, milk vessels, a working shed — and three of four score a perfect 5/5 authenticity. Qwen-Image-2512 is marked down slightly to 4/5: the judge's rationale is specific rather than a vibe, flagging the milking pail as "a modern metal cup with a handle… more typical of Western or urban dairy setups rather than traditional Indian hand-milking practices." One anachronistic object, correctly named, cost exactly one point.

A complete scoring breakdown — one image, every scoring track

Prompt:

Generate a photorealistic photograph of a small vegetable shop in a town in northern India, with a painted signboard above the entrance displaying the Hindi text: सब्जी की दुकान
Z-Image-Turbo generated vegetable shop with an exact-match Hindi sign .9998 .8503 1.000
IndicPhotoOCRdetects, then reads, word by word
सब्जी.9998 की.8503 दुकान1.000
→सब्जी की दुकान
exact match ✓ · CER 0.000 · 3/3 detected boxes assembled left to right
Qwen3-VL-32Bjudges cultural authenticity
5 / 5
“A small vegetable shop in northern India with authentic details: a weathered blue painted signboard with Hindi script, a corrugated metal awning, and fresh vegetables displayed on metal trays … all consistent with local market practices.”
SigLIP2prompt alignment & attributes
Prompt Alignment 0.256 cosine · 5/5 expected attributes present

This is what the other examples are measured against: Z-Image-Turbo on a shorter, more common phrase (three words, "Vegetable Shop") — every track lands well, all at once, on the same image. Across all nine models and the full 300-prompt text-rendering set (English included), only 299 image/model pairs land an exact OCR match out of 2,700 attempted — the bookshop example above, not this one, is the far more common outcome.

Findings

01

One of the nine models tested breaks the wall — and only for one script family.

Z-Image-Turbo is the exception among the nine models scored here: 68.0% exact match on Hindi, 54.0% on Marathi (both Devanagari), a much weaker 10.0% on Bengali, and a flat 0.0% on both Telugu and Kannada — the two Dravidian scripts tested. Its capability is specific to one script family, not a general Indic solution, and even inside that family it is partial.

02

English text-rendering skill does not transfer to Indic scripts.

Qwen-Image-2512 renders English at 96.0% and FLUX.2 [dev] at 92.0% — the two best in the field. Both score exactly 0.0% on every Indic script, identical to Lumina-Image 2.0 and SANA 1.5, whose English is a mere 10.0%. Eight of nine models land at a flat 0% on Devanagari, Bengali, Telugu and Kannada regardless of where their English score falls, which is what a separately trained capability looks like rather than one general "text-rendering quality" dial.

03

0% exact match can still mean the script came out legible — just not perfect.

On the shared example earlier in this report, FLUX.2 [dev] rendered the requested "पुस्तक भंडार" as "पुत्तक भंडार" — one wrong glyph out of twelve (CER 0.083), the kind of slip a reader would likely still parse. SD 3.5 Large's reading of the same sign (CER 0.833) is not Devanagari-shaped text a reader could parse at all. Exact match alone scores both a flat 0% and cannot tell them apart; CER shows the first is a model that can render the script and mostly did, just imperfectly, while the second genuinely cannot.

04

Hypothesis: training data is rich in Latin and CJK text, thin on Indic scripts.

This benchmark scores outputs, not training corpora, so this is an explanation rather than a measurement. Web-scraped caption corpora are dominated by English captions and Latin-script signage, and the models that render Chinese or Japanese reliably got there through corpora deliberately assembled for it. No comparable public corpus pairs images with legible Devanagari, Bengali, Telugu or Kannada text — which would mean a model fails these scripts not from weaker text-rendering skill, but from rarely having seen one.

05

How a model is prompted is itself a variable in any generation benchmark — so holding it fixed is what makes model-to-model comparison meaningful.

Wording, instruction language and examples can all move a generative model's output. This benchmark fixes one sentence template across all 500 prompts and all nine models specifically so a score gap reflects a difference between models rather than a difference in how each one happened to be asked — the same prompt set can be reused as-is to test a tenth model against the exact same bar.

methodology

$ alignment  →  google/siglip2-so400m-patch14-384, checkpoint pinned and recorded per run
$ judge  →  Qwen/Qwen3-VL-32B-Instruct, one call per image, temperature 0
$ ocr  →  Bhashini-IITJ/IndicPhotoOCR — TextBPN++ detection, per-language recognisers
$ normalisation  →  NFC, casefolded, punctuation and zero-width stripped; combining marks (matras) never touched
$ sampling  →  one image per prompt, deterministic seed derived from prompt id, each model run at its own documented default step count and guidance scale — held constant is resolution (1024×1024 for all nine); not held constant is step count, which spans 4 (Z-Image-Turbo) to 50 (FLUX.2 [dev], Qwen-Image-2512, HiDream-I1-Dev), so this is a comparison of each model's own recommended settings, not of rendering skill under a matched compute budget
$ provenance  →  every run's exact checkpoint revision, resolved commit hash, inference settings, dataset SHA-256 and package versions are recorded in that run's run_metadata.json, written before generation starts
$ model set  →  nine open-weight checkpoints, generation and evaluation kept as separate stages so one run can be re-scored without re-generating

credit & scope

None of the scored checkpoints are ours. Every model above is an unmodified public release, run at its own documented defaults. This report contributes the prompt set, the scoring pipeline, and the nine runs — the dataset, the per-run manifests (pinned revisions, config, seeds) and the scoring code that produced every number here ship alongside it in the same repository.