// the fork, in one line
Size on disk is roughly parameters × bits per weight. So at a fixed download size you can spend the budget on more parameters or on more precision per parameter:
That is the nominal figure. Real files run a little bigger: in the llama.cpp quantize README, Q4_K_M is 4.8944 bits per weight on Llama-3.1-8B (4.58 GiB), Q8_0 is 8.5008 (7.95 GiB) and F16 is 14.96 GiB. And the weights are not the whole bill: the context cache sits on top of them and grows with your window (see the KV-cache section of quantization explained). "If it fits" means the file plus the context you actually use.
The measured answer: the bigger one, usually
A 2025 third-party study of Qwen3 quantization (Zheng et al., Beihang / Xidian / ETH Zürich, arXiv 2505.02214; not the Qwen team) quantized every Qwen3 size with the same methods, weight-only, group size 128, and scored them on 5-shot MMLU. Table 4 gives the same-budget pair directly:
The study also reports that bigger models are hurt less by the same squeeze:
A note for anyone checking the tables. Those two percentages match the paper's Table 3 (base models: 14B-Base 80.7 → 79.8, −1.1%; 0.6B-Base 52.3 → 47.2, −9.8%), not Table 4, which the sentence cites. In Table 4 (post-trained models) the same squeeze costs 14B 78.5 → 77.4 (−1.4%) and 0.6B 47.1 → 44.0 (−6.6%; −10.6% with AWQ). The direction holds in both tables; the video draws the base-model numbers and labels them that way.
This "big models lose less" finding is the Qwen3 study's. It is not what the older Dettmers paper below found: its scaling curves for different precisions are "almost parallel".
Against 8 bits, it depends on the method
The honest edge case is the half-size model at eight bits, which is the same download size as the bigger model at four. There the evidence is narrow:
The broad evidence for the rule is older. Dettmers & Zettlemoyer (The case for 4-bit precision, ICML 2023) ran "more than 35,000 experiments ... for 3 to 8-bit precision at scales of 19M to 176B parameters across the LLM families BLOOM, OPT, NeoX/Pythia, and GPT-2" and concluded that "4-bit precision is almost universally optimal for total model bits and zero-shot accuracy." Their recommendation: "keep the precision in 4-bit but vary the number of parameters of the model instead."
Their own abstract frames the question with a posed example, not a measured pair: "a 30B 8-bit model and a 60B 4-bit model have the same number of bits but may have very different zero-shot accuracies." And they name the exception: if the next bigger model does not fit at 4 bits, spend the spare room on bits — "a 48 GB GPU has enough memory to use a 66B model in 5-bit precision but cannot fit a 175B model in 4-bit. Therefore ... 5-bit precision and a 66B model is preferable for this scenario."
The model makers' own numbers agree in direction. The Qwen team's (dated, Qwen2) benchmark page lists Qwen2-72B-Instruct at BF16 81.3, GPTQ-Int8 80.7 and GPTQ-Int4 81.2 (average), while Qwen2-7B-Instruct drops from 66.9 to 64.1 at Int4 — the big model barely moves, the small one does.
Four bits is the floor
The trade stops paying below about four bits. Dettmers & Zettlemoyer's Figure 1 (OPT models): "Zero-shot performance increases steadily for fixed model bits as we reduce the quantization precision from 16 to 4 bits. At 3-bits, this relationship reverses, making 4-bit precision optimal." And: "Pythia and OPT are unstable for 3-bit inference where performance is close to random (35%) for the largest Pythia/OPT models", with one exception, "BLOOM-176B where 3-bit is slightly but not significantly better."
Newer models are not safer down there. The 2025 study: "while Qwen3 maintains competitive performance at higher bit-widths (4-bit and above), it exhibits more pronounced performance degradation compared to previous model generations when quantized to 3-bit or below", and the drop is worst "in complex reasoning tasks and few-shot learning scenarios." (Qwen3 did not fall to random; it lost more than older generations did.)
Q3_K_M is 3.9960 bits per weight. Below four, measure on your own task before you trust it.Two different questions, two different rules
This does not contradict quantization explained. That page's rule is about one model: for the same weights, eight bits is a nearly free drop and four is a real one. This page is about choosing between models at a fixed size, where the bigger model's advantage outweighs what the squeeze costs it.
Where this breaks. Compare within one model family and generation: a newer small model can beat an older big one outright. Quantizers differ (GPTQ and AWQ disagree above). Scores here are MMLU and zero-shot accuracy, not "smartness", and low-bit losses are worst on complex reasoning. And one scaling-law study (Kumar et al., arXiv 2411.04330) predicts that "the degradation introduced by post-training quantization increases as models are trained on more data" — a prediction fitted on models up to 1.7B parameters, but a reason to re-check this rule on the newest, most heavily trained models.
Sources: An Empirical Study of Qwen3 Quantization (arXiv 2505.02214, Tables 3-4, §2.2, §3) · Dettmers & Zettlemoyer, The case for 4-bit precision: k-bit Inference Scaling Laws (arXiv 2212.09720; PMLR 202:7750-7774, ICML 2023) · Qwen docs, Performance of Quantized Models (Qwen2) · llama.cpp tools/quantize README · Kumar et al., Scaling Laws for Precision (arXiv 2411.04330). All verified 2026-09-26.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
