8 MODELS · 1,319 PHOTOGRAPHS · 13 LANGUAGE LABELS
~/research/india-in-the-wild

Indic Script Reading Gap report

8 open-weight VLMsBSTD scene text1,319 photographs

A model gets one full, unedited photograph taken somewhere in India and has to find and read every legible sign in it, in the original script — no crop, no hint about where the text is. Eight open-weight vision-language models, every image in BSTD's official test split, gold answers straight from the dataset's own annotations. The best F1 in the suite is 0.4715, and the top two models are separated by 0.0055 on the fuzzy variant and tied on the exact one. The decisive gap is not between models but between scripts: on the 914 images that carry both, every model reads the English better than the Indic text beside it — the strongest of them 0.79 against 0.21.

report · BSTD scene-text dataset · 8 models, 13 language labels · Aug 2026
~/research/india-in-the-wild

Indic Script Reading Gap

Eight open-weight vision-language models, attempted on all 1,319 photographs in BSTD's official test split · Aug 2026

the task

This benchmark asks a vision-language model to read Indian public signage the way it is actually photographed: a hand-painted shop board at an angle, a fare list behind a grille, a hoarding half in shadow, multiple languages on one signboard. Each model is shown one full, unedited test-set image from BSTD — a scene-text dataset of real photographs from across India, annotated at word level with a polygon, the exact transcription, and a script-language label — and asked to find and read everything legible in it.

Every model is given the identical prompt, verbatim:

You are looking at a real, unedited photograph taken somewhere in India. Read every piece of legible text visible in the image: shop signs, boards, banners, posters, hoardings, labels, name plates, price lists, anything with writing on it. Transcribe each distinct piece of text exactly as it is written, in its original script. Do not translate or transliterate. Do not describe the image or the objects in it. List one piece of text per line, in the order you notice them. If a piece of text is present but not legible, skip it rather than guessing. Output only the transcribed lines, nothing else.

No bounding boxes are given, no script is named and no example is shown, so the model has to both find the text and read it.

Because BSTD supplies the gold strings, scoring is a string comparison rather than a model-judged rubric: the sign says what it says.

metrics

recall
Of the strings BSTD's annotators marked, what share did the model produce? Each gold string is checked once against the model's whole output. Text the annotators skipped cannot lower it.
F1
Harmonic mean of recall and precision, so a model has to do both: transcribe most of what is on the sign, and little that isn't. Its value depends on the matching rule.
precision
Of the lines the model produced, what share is backed by some gold string? BSTD's annotators did not mark every legible word, so a model that correctly reads an unannotated sign is counted wrong for it — this is a floor, not an estimate. Matching is per output line, so how a model breaks its lines moves the number too.
exact vs. fuzzy
exact requires the gold string as a normalized substring; fuzzy also accepts a match at 0.8 similarity (difflib.SequenceMatcher), forgiving a dropped matra or one swapped character but not a wrong reading. The leaderboard gives both; the per-script tables give fuzzy.

Findings

01

Every model reads English better than the Indic script beside it.

All eight, without exception, on the 914 test images that carry both. The gap runs from 1.4× at the narrowest to 23× at the widest — the strongest model recalls 0.79 of the English text against 0.21 of the Indic text in the same photographs. Camera, exposure and framing are constant by construction, though text size and contrast can still differ within a frame.

02

Without an instruction to stay silent, no model does.

58 photographs carry no legible text at all by BSTD's own annotators. Across all eight models — 464 model-image pairs — the abstention count is zero. The prompt says to skip an illegible word, but never says what to output when nothing in the frame is legible, so this measures a gap in the prompt as much as a property of the models.

03

4.6× the parameters, and the F1 is a wash.

Qwen2.5-VL-32B scores 0.4715 fuzzy F1 against the 7B's 0.4660 — and on the exact variant it is fractionally behind, 0.4527 to 0.4530. A cluster bootstrap over images puts the gap at +0.0055 with a 95% interval of [−0.031, +0.041] — the two are not separable, and the 7B comes out ahead in 39% of resamples. What the size buys is output volume: 42,036 lines against 19,406, which lifts recall by 0.063 and costs 0.080 of precision.

04

Digits survive where the script around them does not.

Split the gold strings into purely numeric ones (1,082 of them — phone numbers, prices, distances, years) and Indic-script ones, and every run recalls the digits several times better: Qwen-32B 0.680 against 0.212, Gemma-3-27B 0.347 against 0.195, MiniCPM-V-2.6 0.214 against 0.013. Numerals are shared with Latin script; the letters beside them are not. They are not simply the easy end of the English side either: on the same signs the ordering for all eight models is Latin letters, then numerals, then Indic script.

instances of the dataset

BSTD is a scene-text dataset of 6,582 photographs taken across India — 5,263 in its train split and the 1,319 of its test split scored here — each annotated by hand at word level: a polygon around every piece of text, the exact string inside that polygon, and the language it is written in. Those annotations are the ground truth; a transcription is graded by string comparison against them, with no model judging the answer.

Every item in this benchmark looks like the three below: one full, unedited photograph, and a set of human-drawn polygons each carrying the exact string inside it and the language it is written in. Nothing is cropped for the model and nothing is straightened — it is handed the whole frame and asked to transcribe every piece of legible text it can find, in the original script, one per line. The yellow outlines below are BSTD's ground truth, not model output.

Why bilingual signage is the useful case: India's public signage is routinely multilingual, so a single photograph often carries English and an Indic script on the same board, in the same paint, at the same distance from the lens. That makes each such image a partial control: if a model reads the Latin line and invents the Devanagari one beside it, the camera, the lighting and the framing were the same for both. 914 of the 1,319 test images are bilingual in this way.

Doraha station board Punjab

A railway station board with three scripts stacked on one panel: Gurmukhi on top, Devanagari and Latin below, and an elevation line beneath them in small Devanagari and Latin digits.

The full photograph of the Doraha station board, with BSTD ground-truth polygons outlined in yellow.the whole frame, exactly as the model receives it
ground truth 9 annotations
  • DORAHAenglish
  • दोराहाhindi
  • ਦੋਰਾਹਾpunjabi
  • समुद्रhindi
  • तलhindi
  • सेhindi
  • ऊंचाईhindi
  • 258.92english
  • मीटरhindi
what came back every line the model produced, verbatim; green = backed by an annotation
Qwen2.5-VL-32B1/3 supported
रेलवे
दूराहा DORAHА
स्टेशन तल से 258.92 मीटर
Qwen2.5-VL-7B2/2 supported
दोराहा DORAH
स्वमंडु तल से ५ चार्डे 258.92 मीटर
Gemma-3-27B1/3 supported
ਟੇਹਲਾ
ਡੋਰਾਹਾ DORAHA
ਲੁਧਿਆਣਾ ਜੰ: ਤੋਂ 6 ਕਿ:ਮੀ: 256.92 ਮੀਟਰ

Kotilingala temple plaque Telangana

A weathered bilingual plaque under a 13th-century Vishnu relief. The Telugu and the English say the same thing, side by side, in the same enamel and the same state of wear.

The full photograph of the Kotilingala temple plaque, with BSTD ground-truth polygons outlined in yellow.the whole frame, exactly as the model receives it
ground truth 9 annotations
  • కోటిలి౦గాలtelugu
  • 13వtelugu
  • క్రీ॥శ॥telugu
  • VISHNUenglish
  • విఘ్ణtelugu
  • శతాబ్థముtelugu
  • 13thenglish
  • C.A.D.english
  • Kotllingalaenglish
what came back
Qwen2.5-VL-32B3/6 supported
VISHNU
13th C.A.D.
Kotilingala
మూలం
శ్రీకృష్ణా కుటుంబం
కొత్తింగలం
Qwen2.5-VL-7B4/6 supported
విష్ణు
క్రీ.శ 13వ శతాబ్దం
కొట్టింగల
VISHNU
13th C.A.D.
Kotilingala
Gemma-3-27B5/6 supported
విష్ణు
శ్రీశైలం 13వ శతాబ్దము
కొటిలింగాల
VISHNU
13th C.A.D.
Kotlingala

Squeeze cage exhibit label Karnataka

A museum label on a bamboo post beside a wheeled iron cage. Two lines on one small board — four Kannada words above, the four English words that translate them below — at the same angle, the same distance from the lens and the same exposure.

The full photograph of the Squeeze cage exhibit label, with BSTD ground-truth polygons outlined in yellow.the whole frame, exactly as the model receives it
ground truth 8 annotations
  • ಪ೦ಜರkannada
  • Cageenglish
  • Squeezeenglish
  • Royalenglish
  • Heritageenglish
  • ರಾಜರkannada
  • ಕಾಲದkannada
  • ಸ್ಕ್ವೀಜ಼್kannada
what came back
Qwen2.5-VL-32B1/2 supported
राजा राणी नौका मोड़
Royal Heritage Squeeze Cage
Qwen2.5-VL-7B1/2 supported
ರಾಜ್ಯ ಕಲ್ಲು ಮುಂಭಾಗದ ಮಂಡಪ
Royal Heritage Squeeze Cage
Gemma-3-27B1/2 supported
ರಾಜ ರಕ್ತದ ಸಿಂಹ ಸೊಂಡೆ ಚಂದ್ರ
Royal Heritage Squeeze Cage

the same-image control

Fuzzy recall on the 9,512 English and 11,785 Indic-script annotations that appear on the same 914 images. Each pair of dots is one model; the bar between them is that model's script gap, with image quality held constant by construction.

Of the 1,319 test images, 914 carry both at least one English annotation and at least one Indic-script annotation.

English annotations Indic-script annotations
0.00.20.40.60.8fuzzy recall on the same 914 imagesQwen2.5-VL-32BQwen2.5-VL-32B — Indic scripts 0.211Qwen2.5-VL-32B — English 0.7890.2110.789Qwen2.5-VL-7BQwen2.5-VL-7B — Indic scripts 0.230Qwen2.5-VL-7B — English 0.6140.2300.614Gemma-3-27BGemma-3-27B — Indic scripts 0.187Gemma-3-27B — English 0.4320.1870.432MiniCPM-V-2.6MiniCPM-V-2.6 — Indic scripts 0.014MiniCPM-V-2.6 — English 0.3260.0140.326AriaAria — Indic scripts 0.071Aria — English 0.3170.0710.317LLaVA-1.6-34BLLaVA-1.6-34B — Indic scripts 0.016LLaVA-1.6-34B — English 0.2360.0160.236Pixtral-12BPixtral-12B — Indic scripts 0.044Pixtral-12B — English 0.2140.0440.214Aya-Vision-8BAya-Vision-8B — Indic scripts 0.050Aya-Vision-8B — English 0.0690.0500.069

leaderboard: every run, overall

BSTD's full test split, one run per model, less the one or two images per model that failed every retry. exact is a normalized substring match; fuzzy additionally accepts a match at 0.8 similarity, so it is always the higher of the pair.

#runrecallprecisionF1
exactfuzzyexactfuzzyexactfuzzy
01Qwen2.5-VL-32B
Qwen/Qwen2.5-VL-32B-Instruct
0.4440.4600.4610.4830.4530.471
02Qwen2.5-VL-7B
Qwen/Qwen2.5-VL-7B-Instruct
0.3860.3970.5480.5640.4530.466
03Gemma-3-27B
google/gemma-3-27b-it
0.2920.3020.4140.4340.3430.356
04Aria
rhymes-ai/Aria
0.1760.1790.2420.2450.2040.207
05MiniCPM-V-2.6
openbmb/MiniCPM-V-2_6
0.1510.1540.3760.3950.2160.221
06Pixtral-12B
mistralai/Pixtral-12B-2409
0.1190.1210.3220.3300.1730.178
07LLaVA-1.6-34B
llava-hf/llava-v1.6-34b-hf
0.1150.1160.4390.4420.1820.184
08Aya-Vision-8B
CohereForAI/aya-vision-8b
0.0560.0590.2970.3080.0940.098
Decoding is identical across rows except for one setting: Pixtral-12B and LLaVA-1.6-34B ran at repetition_penalty 1.30, the other six at 1.05, so any comparison involving those two rows is not strictly controlled — neither low value was a real option, though: at 1.05, Pixtral fell into repetition loops on 21.7% of images, and LLaVA at vLLM's own default degenerated on effectively all of them.

results by dominant language

One table per metric, because the three answer different questions: a model can score well on precision by writing very little, and that looks nothing like scoring well on recall by reading accurately. Rows are the eight models. Columns group the images by their dominant language label — the most common one among that image's annotations. These are BSTD's own labels rather than scripts: Hindi and Marathi share Devanagari, Assamese and Bengali share Eastern Nagari, so the thirteen columns span eleven writing systems. All three tables are the fuzzy variant, the more generous of the two matchers; the exact variant is in the leaderboard above. An outlined cell is the best model in that column, except in Meitei — at 130 gold strings, less than half the next-smallest column, a few hits move it enough that a "best" badge there would mark a coin flip, not a result, so the outline is suppressed for that column only.

recall / precision / F1<.05.05.15.30.45.60+

recall, fuzzy — of the annotated text, how much was found

modelall imagesassamesebengalienglishgujaratihindikannadamalayalammarathimeiteiodiapunjabitamiltelugu
Qwen2.5-VL-32B0.4600.3040.4450.6130.1540.5170.1930.1920.4390.3690.0980.1300.2570.172
Qwen2.5-VL-7B0.3970.2530.4170.5080.1470.5060.1560.1790.4190.2920.0960.1070.2980.122
Gemma-3-27B0.3020.2370.1990.3720.1420.3580.1610.1130.2690.2850.1310.1700.3130.196
Aria0.1790.0520.0480.2600.0310.2300.0690.0480.0920.2080.0590.0460.0680.058
MiniCPM-V-2.60.1540.0760.0350.2450.0280.1060.0530.0480.0370.0460.0550.0360.0560.079
Pixtral-12B0.1210.0520.0380.1800.0070.1330.0480.0310.0530.0080.0370.0450.0650.032
LLaVA-1.6-34B0.1160.0580.0160.1820.0160.1110.0210.0170.0340.0230.0210.0340.0320.034
Aya-Vision-8B0.0590.0150.0250.0740.0110.0960.0160.0220.0430.1690.0060.0330.0290.026

precision, fuzzy — of what the model wrote, how much is backed by an annotation

modelall imagesassamesebengalienglishgujaratihindikannadamalayalammarathimeiteiodiapunjabitamiltelugu
Qwen2.5-VL-32B0.4830.5420.3360.5730.4030.4800.4460.1990.3630.2810.3110.1600.2280.722
Qwen2.5-VL-7B0.5640.4740.6000.7350.3440.5830.3390.4170.7060.8670.1550.3990.4930.088
Gemma-3-27B0.4340.2490.4140.5770.2010.4000.2570.4460.5550.7650.3060.4370.8080.089
Aria0.2450.2220.0890.2950.1160.2810.1950.0800.3570.8660.2280.1240.0510.050
MiniCPM-V-2.60.3950.5490.1730.5610.0970.3680.0460.2830.4770.0070.1170.1390.0240.212
Pixtral-12B0.3300.2430.1880.5010.0320.4460.1170.0660.4000.0620.0900.3290.1150.022
LLaVA-1.6-34B0.4420.2400.1690.5510.0680.6160.0410.0240.3410.3330.0820.3050.3620.085
Aya-Vision-8B0.3080.5740.7270.3040.0490.4940.0340.0550.3480.9850.0130.4190.0490.100

F1, fuzzy — the two combined

modelall imagesassamesebengalienglishgujaratihindikannadamalayalammarathimeiteiodiapunjabitamiltelugu
Qwen2.5-VL-32B0.4710.3900.3830.5920.2230.4970.2700.1950.3980.3190.1480.1430.2410.278
Qwen2.5-VL-7B0.4660.3300.4920.6010.2060.5420.2140.2510.5250.4370.1190.1690.3710.102
Gemma-3-27B0.3560.2430.2690.4520.1660.3780.1980.1800.3620.4150.1840.2450.4510.123
Aria0.2070.0840.0620.2770.0490.2530.1020.0600.1460.3350.0930.0680.0580.054
MiniCPM-V-2.60.2210.1340.0580.3410.0440.1650.0490.0820.0690.0130.0750.0570.0340.116
Pixtral-12B0.1780.0860.0640.2640.0120.2050.0680.0430.0940.0140.0530.0800.0830.026
LLaVA-1.6-34B0.1840.0940.0300.2740.0260.1880.0280.0190.0610.0430.0340.0610.0590.049
Aya-Vision-8B0.0980.0300.0480.1190.0180.1610.0220.0320.0770.2890.0090.0620.0370.042

methodology

$ dataset  →  BharatSceneTextDataset, word-level polygons carrying the exact string and its language label
$ split  →  every image BSTD marks test — 1,319 images, 24,994 gold strings after cleaning; each model scores 24,987 or 24,988, the 1–2 that failed every retry being excluded rather than zeroed
$ cleaning  →  free-text language labels normalized; 18,346 "UNK"/"NA"-placeholder annotations dropped (14.5% of the raw set), unevenly by language
$ model set  →  eight open-weight checkpoints, chosen for a working vLLM integration
$ serving  →  vLLM, temperature 0, max_tokens 1024, one pass per image
$ scoring  →  NFC-normalized and casefolded; exact = substring, fuzzy = difflib.SequenceMatcher ratio ≥ 0.8; ordering is not checked

credit & links

BSTD is not ours. It is BharatSceneTextDataset, a scene-text dataset of real photographs from across India, annotated at word level across 14 language labels, and every gold string scored on this page is its annotators' work. This report contributes the reading task, the scoring and the eight model runs.