What INT4 quantization actually costs
Most guidance on quantization says roughly the same thing: expect to trade a few points of accuracy for a large saving in memory. Useful, and in aggregate it is true.
It is also the wrong shape of advice, because the damage is not spread evenly. We ran the same INT4 checkpoint across four of our benchmarks and watched it score 4.6% on one and place 5th of 21 on another. Not a model that degraded. A model that failed completely at one kind of task and performed normally at another.
If you are deploying a quantized checkpoint, that difference is the thing worth knowing, and no single benchmark score will show it to you.
// what we measured
Gemma 3 12B, in its standard and INT4 forms, run through four benchmarks with different task shapes.
| Benchmark | What it asks for | Gemma 3 12B | INT4 |
|---|---|---|---|
| BKP-500, India-specific knowledge | An exact value, in a constrained output format | 35.1% (rank 10 of 19) | 4.6% (rank 19 of 19) |
| Indic-KCC, crop advisory | Open-ended advice, judged on four axes | 2.30 (rank 6 of 21) | 2.40 (rank 5 of 21) |
| MILU, multilingual knowledge | Multiple choice | 56.41 | 53.94 |
| Code-mixing, maths word problems | An answer under script variation | 80.4% retention | 77.8% retention |
Three of those four results are unremarkable. MILU costs two and a half points. Code-mixing costs under three. Crop advisory costs nothing at all, and in fact the quantized checkpoint edges ahead, which is within noise but certainly not damage.
The fourth is a collapse.
// what the collapse looks like up close
On BKP-500 the quantized model does not score badly. It largely stops participating.
| Measure | Gemma 3 12B | INT4 |
|---|---|---|
| Bharat Score | 35.1% | 4.6% |
| Refusal rate | 0.0% | 91.3% |
| Unit discipline | 83.2% | 2.0% |
| Out-of-magnitude error rate | 25.6% | 1.4% |
| Consistency | 99.6% | 100.0% |
Two of those numbers look like improvements and are not.
The out-of-magnitude error rate falls from 25.6% to 1.4%. That measures how often a numeric answer is wrong by a factor of ten or more, which on lakh and crore scales is the characteristic failure. It drops because the model is no longer producing numbers to be wrong about.
Consistency rises to a perfect 100%, meaning the model agrees with itself across every repeated sample. It is reliably declining to answer.
The per-category breakdown makes it plain. The quantized model scores a flat zero on five of the seven categories:
| Category | Gemma 3 12B | INT4 |
|---|---|---|
| Indian numeral system | 48.8 | 0.0 |
| Weights, volumes and informal measures | 40.7 | 0.0 |
| Land units by state | 20.8 | 0.0 |
| Agricultural seasons and crop calendars | 41.9 | 0.0 |
| Fiscal-year conventions | 37.8 | 0.0 |
| Government schemes | 29.0 | 12.2 |
| Structural identifiers | 26.7 | 20.0 |
For context, the 4B version of the same model family scores 25.3% overall on this benchmark. The quantized 12B scores 4.6%. You would be better served by a model a third of the size.
// why the two results differ
The four benchmarks differ in one way that tracks the result exactly: what they require of the output.
BKP-500 asks for an exact value under a strict format. Every item has a declared answer type, a gold value, and a tolerance, and grading is done by code rather than by another model. A response that hedges, waffles, or declines is marked incorrect, because for this task it is incorrect.
Indic-KCC asks for open-ended agronomic advice, judged by a model on correctness, naturalness, groundedness and safety. There is no single right string. A response that is approximately right in an unexpected shape still scores.
MILU is multiple choice. Code-mixing maths sits between the two.
So the pattern is that quantization damage in this case is concentrated in constrained, exact-value output, and is close to invisible in open-ended generation. That ordering held across every benchmark we ran.
We would not generalize this to all quantization. We measured one checkpoint, and what we can say is that its failure is task-shaped rather than uniform.
One detail makes the result harder to dismiss. The checkpoint we tested is quantization-aware trained, which is the pathway specifically designed to preserve quality through compression. A naive post-training quantization failing would be unremarkable. A QAT checkpoint scoring zero on five of seven categories is not what that method is supposed to produce.
// what this means if you are deploying
An aggregate benchmark score will not warn you. A quantized model that looks fine on a general leaderboard can be entirely unusable for the thing you actually need, and the leaderboard has no way to show you that.
Test on the output shape your pipeline requires. If you need JSON, or a number with a unit, or a value inside a tolerance, test exactly that. Most production pipelines need constrained output, which is precisely the case that broke here.
Watch refusal rate, not just accuracy. A 91.3% refusal rate is visible immediately and is a far clearer signal than an accuracy number that quietly reflects a model declining most of the work.
Compare against a smaller unquantized model. On this benchmark, the 4B checkpoint beats the quantized 12B by more than five times. If quantizing a large model puts you below a small one, quantization is not buying you anything.
Pin the revision. Quantized mirrors are often community-maintained, and their default branches move. If you have validated a build, pin the commit rather than tracking a branch, or you may be running different weights next month without knowing it. This is advice we are taking ourselves, for reasons set out below.
// method and caveats
All figures come from Sthānika AI's own benchmark runs. Each benchmark's corpus, evaluation harness, and raw model responses are published.
- BKP-500: 552 items across seven categories, graded deterministically by code with no model in the grading loop. Method → · 2026 results →
- Indic-KCC: 500 Kisan Call Centre questions in 11 languages, judged by a reference-grounded model on four axes. Method → · 2026 results →
- MILU 2026 → · Code-mixing tax →
Caveats worth stating. This is one checkpoint at one quantization level. We did not test other quantization methods or other model families, so the result should be read as evidence that quantization damage can be task-shaped, not as a measurement of INT4 in general. The BKP-500 and Indic-KCC rosters are not identical, so the two ranks are not directly comparable to each other. And as with everything we publish, gold answers on BKP-500 are currently visible, so scores should be read alongside a model's training cutoff.
Which checkpoint. The quantized model tested was unsloth/gemma-3-12b-it-qat-int4-bnb-4bit, an ungated mirror of Google's gemma-3-12b-it-qat-int4-unquantized, served through vLLM with --quantization bitsandbytes at temperature 0, three samples per item across both prompt regimes.
A reproducibility gap of our own. This checkpoint was not commit-pinned in our run registry, so it resolved to whatever the mirror's default branch served on the run date. Two of the nineteen models in the roster carry explicit revision pins; this was not one of them. We therefore cannot rule out that one specific build is responsible rather than the quantization approach, and we would welcome anyone reproducing this against a pinned revision. The full configuration, including the vLLM launch arguments and the batch-sizing notes, is published at configs/models/gemma3-12b-int4.yaml.
// how to cite
Sthānika AI (2026). What INT4 quantization actually costs. https://sthanika.ai/notes/int4-quantization-cost
