~/methodology

How we build benchmarks

// why this page exists

Anyone can run a model against a list of questions and publish a table. What makes a benchmark worth citing is everything that happens before the table: where the questions came from, how the answers were verified, what the grader actually checks, and what the authors are willing to admit they got wrong.

This page describes how we do that work, using the two benchmarks we have published as worked examples.

// the starting point: public data is published for its own purpose

India publishes a great deal of useful data. Farmer helpline transcripts, government notifications, RBI and NSO publications, state revenue department references, scheme documentation. Much of it is open, and the institutions that maintain it deserve credit for that.

But data published for administrative record-keeping, public transparency, or service delivery is not the same thing as data structured for machine evaluation, and it was never meant to be. A call-centre transcript exists so a query can be answered and logged. A revenue department circular exists so an officer can apply a rule. Neither exists so a language model can be scored against it.

Converting one into the other is real work, and it is the work we do. The gap is not a criticism of the source. It is a difference in purpose, and closing it is what turns a public record into a research-grade corpus.

That conversion is where errors get introduced if nobody is careful, and it is where most of our effort goes.

// how a corpus gets built

1. Source and record provenance

Every fact enters through a named source, and the source travels with it.

For BKP-500, the numeric and unit facts come from a hand-curated registry drawn from government notifications, RBI and NSO publications, and state revenue department references. Each row carries a provenance field holding the source, URL, supporting quote, and a confidence rating.

For Indic-KCC, the questions come from Kisan Call Centre transcripts, reaching us through ICAR's KCC-CHAKSHU portal and a CC0-licensed public mirror.

Provenance is recorded per row rather than per dataset. That distinction matters: a dataset-level citation tells you roughly where a corpus came from, while a row-level one lets a reader check any individual item we got wrong.

2. Structure the content for evaluation

Raw source material rarely arrives in a form a grader can score. A helpline exchange is a conversation. A circular is prose. Neither has a gold answer attached, a declared answer type, or a tolerance.

So every item is given an explicit structure: the question text, the gold answer, the answer type (numeric, numeric with unit, date, date range, month set, enum, normalized string, or clarification), a grading tolerance where numeric, and any accepted alternate phrasings.

This is the step where a benchmark either becomes reproducible or does not. A gold answer without a declared type and tolerance cannot be graded consistently by anyone but its author.

3. Test conventions, not trivia

Most "India knowledge" questions test trivia, which frontier models already handle well and which leaks into training data quickly. We design items around something harder to fake: the conventions a resident considers too obvious to mention.

How many square feet is a kattha in Muzaffarpur? Which quarter of the Indian fiscal year does November fall in? What is "paune crore" as a plain number of rupees? Nobody in India would think these were difficult. In our 2026 run, ten items of exactly this kind were answered correctly by none of the 19 models tested.

Where the right answer depends on the jurisdiction, the item records which state it applies to. Where official sources conflict, the item is built to reward recognizing the ambiguity rather than committing to a single number.

4. Translate, where translation is required

Indic-KCC presents all 500 questions in 11 languages: English plus Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam, Punjabi and Odia. That is 5,500 prompts, and every one passes through our own translation and quality-control pipeline.

This is where a benchmark can quietly break. A translation error does not announce itself as an error; it shows up as a model scoring badly in one language, and it is indistinguishable from a genuine capability gap unless the pipeline is checked.

5. Review, then review again

Every item goes through human review against its recorded source. Items where the reviewer disagrees with the gold answer are flagged rather than silently corrected.

Where a flagged item reflects a convention that conflicts across sources, we leave it open rather than resolving it mechanically. A state land unit with two defensible definitions in two different official documents is not a data-entry error to be fixed. It is a real ambiguity, and pretending otherwise would make the benchmark look cleaner than the world is.

For BKP-500, all 552 items have been through a two-reviewer pass. The reviewers flagged 57 items with a problem in the gold answer; 41 were corrected, and 16 were kept as clarification items because the sources themselves disagree. Corpus-level adjudication is complete for all 552.

6. Grade with code, not with a model

BKP-500 is graded entirely deterministically. Code decides right and wrong, using one of eight graders selected per item from its declared answer type. There is no language model in the grading loop.

When a grader cannot extract a comparable answer from a response, it declines rather than guessing, and that response is recorded as undetermined. It counts in neither the correct nor the incorrect column.

Indic-KCC cannot be graded this way, because open-ended agronomic advice has no single correct string. It uses a reference-grounded language model judge scoring four axes separately, and we publish the parse rate alongside every score so a reader can tell a poor answer from an unparseable one.

The choice between the two approaches is driven by the task, and we say which is in use on every benchmark page.

7. Protect the benchmark from contamination

Every BKP-500 item carries a unique canary string. Anyone assembling a pretraining corpus who wants to honor benchmark-exclusion norms can filter on it.

This costs us nothing and protects the benchmark's usefulness for everyone who comes after. We would encourage it as a default for any evaluation corpus published openly.

8. Change the benchmark when the evidence says to

The first published version of BKP-500 paired every item with an internationally framed control question, and reported the gap between the two as a headline metric. The idea was sound: a control holds general ability constant, so the difference should isolate India-specific knowledge.

Running it at scale showed that the control set had systematic design flaws. The metric built on it could not carry the weight we had put on it.

So we removed both. The corpus went from 1,086 items to the 552 India-specific ones, every score was recomputed, and the change is recorded on the benchmark page with a date rather than edited in quietly.

We include this here because it is the clearest evidence we can offer for everything else on this page. A benchmark that has never been corrected has usually never been checked.

// what we publish

Everything needed to check our work, or to disagree with it:

ArtifactWhere
The corpus, with per-row provenanceHugging Face
The evaluation harness, adapters, graders and scoringGitHub
Raw per-model responses and exact run configurationGitHub
Method, metric formulas and known limitsBKP-500 · Indic-KCC

A score you cannot reproduce is an assertion, not a measurement. Publishing the run configuration alongside the results is what makes the difference.

// known limits, stated up front

Every benchmark we publish carries its own limitations section. A few that apply today:

We would rather publish a benchmark with its limits named than a cleaner-looking one with them omitted. A reader who finds a weakness we did not disclose has every reason to doubt the rest.

// standing commitments

We exclude our own models. Every checkpoint we have fine-tuned is excluded from every benchmark roster we publish. If our models are ever reported against our own benchmarks, they will be labeled as ours and reported separately.

We correct things in public. When we get something wrong, the correction goes on the page, dated, rather than being quietly edited in.

We do not accept payment to evaluate a model. No organization has paid for inclusion, placement, or the timing of a result.

// working with us

If you maintain a dataset that could support an evaluation corpus, or you build models you would like included in a future run, get in touch. We are particularly interested in working with institutions that already publish open data, because the corpus is more useful to everyone when the source maintainers are part of building it.

Last updated: September 22, 2026