~/glossary/locale-gap-delta

What is Locale Gap Delta?

In one line: a model's accuracy on international control items minus its accuracy on the matched India-specific items they are paired with.

// definition

Locale Gap Delta is a model's accuracy on international control items minus its accuracy on the matched India-specific items they are paired with. It is computed over BKP-500 matched pairs only, and reported in percentage points.

Locale Gap Delta = accuracy(control items) − accuracy(India items)

// how to read it

Positive means the internationally framed version of a question was easier for the model than its India-framed twin. This is the expected direction for a model with no India specialization, and it is what the metric was built to detect.

Negative means the India-framed item was easier. This does happen, and among the stronger models in the BKP-500 roster it is common rather than exceptional.

Near zero means the two framings were of similar difficulty for that model.

// why the pairing matters

A low score on India-specific items, taken alone, tells you very little. It could mean the model does not know Indian conventions, or it could mean the model is weak at the underlying arithmetic, weak at date handling, or poor at following the required answer format.

Every India item in BKP-500 is therefore matched to a control item that asks the same underlying question under an international or self-contained convention, at comparable difficulty. The control holds general capability constant. When a model handles the control and fails its paired India item, the difference is attributable to the India-specific convention rather than to general weakness.

Eighteen core items are deliberately unpaired. These are clarification-type items with no international twin, where the correct behavior is to hedge rather than commit, and they are scored on Ambiguity Handling and Overconfidence instead.

// one caution when citing it

An extreme Locale Gap Delta from a model with a very low Bharat Score is usually a symptom of broad failure, not a clean locale signal. In the 2026 run the largest positive value in the roster, +29.6%, belongs to a quantized checkpoint that scores zero on five of the seven categories and refuses nearly half of all items. At that failure rate the paired comparison stops carrying its intended meaning.

The metric is most informative on models that are otherwise working.

// a finding worth knowing

In the 2026 BKP-500 run, the most negative Locale Gap Delta values belong to the strongest models overall, several of them global rather than India-built. The measure correlates with general capability, so it should be read alongside Bharat Score rather than as an independent claim about India specialization.

// related

Bharat Score is the companion metric, measuring macro-mean accuracy across the seven BKP-500 categories.

// where the term comes from

Locale Gap Delta was defined by Sthānika AI for the BKP-500 benchmark.

// citation

Sthānika AI (2026). BKP-500: the Bharat Knowledge Probe. https://sthanika.ai/benchmarks/bkp-500