~/about

Introducing Sthānika AI

·

A model can write fluent Hindi and still think a bigha is the same size in Punjab and Bihar. It can discuss Indian agriculture at length and put the kharif season in the wrong months. It can convert ₹1.2 lakh crore into dollars with total confidence and be wrong by a factor of a hundred.

None of those are language failures. They are the failures of systems built somewhere else, trained mostly on text from somewhere else, and then asked to work here.

Sthānika AI exists to measure that gap precisely, and then to close parts of it.

Sthānika means "of this place."

// What we do

We are an independent research lab in Hyderabad. We do two things.

We build benchmarks. Public Indian data is published for administrative purposes: helpline transcripts, government notifications, revenue department circulars, scheme documentation. It is not structured for evaluating machine learning models, and was never meant to be. Turning one into the other is the work. We source with per-item provenance, structure for deterministic grading, translate and review, then publish the corpus, the harness and the raw model responses so anyone can check our numbers or disagree with them.

We build models. Small, specialist, open. Fine-tuned for the tasks where general-purpose models are structurally weak, and small enough to run on a single machine or inside an institution's own infrastructure.

// Why both

Because measurement has to come first, and almost nobody does it that way.

No model of ours trains before its evaluation set exists. That ordering is the whole discipline: it is what stops a team optimizing for a benchmark it wrote after the fact, and it is why every claim on this site traces back to a published measurement rather than an impression.

It also means our benchmarks are useful independently of our models. Most of the systems we measure are not ours and never will be.

// What we have published

Since June 2026:

Eight benchmark studies, spanning language, speech and vision. 19 models on India-specific knowledge. 21 on crop advisory across 11 languages. 17 on multilingual accuracy, run on the complete test set with no subsampling. 16 speech-to-text models across 22 languages and more than 357,000 utterances. 16 on code-mixed and romanized input. Nine text-to-image models on Indic script rendering. Eight vision-language models reading signage in 1,319 unedited photographs. Thirteen tokenizers measured for what Indian languages actually cost to process.

Two benchmarks we maintain, with their corpora, graders and metric definitions published: BKP-500 for India-specific knowledge, and Indic-KCC for crop advisory.

Two model families, released open with weights, training data and evaluations: Sieve, a family of decision models returning calibrated probabilities rather than generated text, and an Indic crop-advisory adapter that outperforms a base model with more than twice its parameters.

Every dataset is on Hugging Face. Every harness is on GitHub.

// How we work

Specialist beats general. A small model fine-tuned on the right data outperforms frontier models on its task, at a fraction of the cost. We target where frontier models are structurally weak: code-mixed Indian languages, India-specific domains, and deployments that have to run on-premise.

Open is the strategy. Every project ships its weights, its curated dataset and its expert-graded benchmark. Reproducibility is the product. A score you cannot reproduce is an assertion, not a measurement.

Evaluation first. No model trains before its evaluation set exists.

We exclude our own models from our own benchmarks. Every checkpoint we have fine-tuned is left out of every roster we publish. If our models are ever reported against our own benchmarks, they are labeled as ours and reported separately.

We correct in public. When we get something wrong, the correction goes on the page with a date rather than being edited in quietly. There is a worked example on the BKP-500 benchmark page, where we removed part of the benchmark after evaluation showed it did not hold up.

How we build our benchmarks →

// The partner covenant

We build with partners who hold deep domain knowledge: seed companies, hospitals, diagnostics chains, public institutions. The arrangement is the same every time.

The general capability is released open. The partner's proprietary edge stays theirs.

An open front door and a private deep end. It is the only structure we have found that lets a lab publish without constraint while doing work that is commercially worth doing.

// Who we are

Sthānika AI is the research lab of PurpleTalk, a twenty-year digital innovation company headquartered in Hyderabad. Where PurpleTalk builds products, Sthānika publishes research.

The lab was announced in June 2026.

// Working with us

Partners. You hold the domain knowledge and the data exhaust; we build the model. Propose a partnership →

Researchers and engineers. Small team, real GPUs, everything you build gets published. See open roles →

Institutions and government. Sovereign, on-premise, open-weight deployments, evaluated in the open. Start a conversation →

// The aim

To be the lab India's institutions trust with the models that run closest to their people.