Qwen3 8B quants compared
24 GGUF builds by real file size, probed from bartowski/Qwen_Qwen3-8B-GGUF on Hugging Face (2026-09-02). 8B params.
Download Qwen3 8B Q4_K_M (4.68 GB) — it fits 8 GB of memory with 16k context. With Ollama: ollama run qwen3:8b-q4_K_M
Every Qwen3 8B quant by real file size
| Quant | Weights | + KV (16k) | Total | Fits comfortably in | Quality |
|---|---|---|---|---|---|
| BF16 | 15.26 GB | 2.0 GB | 17.3 GB | 24 GB | Full precision (lossless) |
| Q8_0 | 8.11 GB | 2.0 GB | 10.1 GB | 12 GB | Near-lossless |
| Q6_K_L | 6.54 GB | 2.0 GB | 8.5 GB | 12 GB | Excellent |
| Q6_K | 6.26 GB | 2.0 GB | 8.3 GB | 12 GB | Excellent |
| Q5_K_L | 5.81 GB | 2.0 GB | 7.8 GB | 12 GB | Very high |
| Q5_K_M | 5.45 GB | 2.0 GB | 7.5 GB | 12 GB | Very high |
| Q5_K_S | 5.33 GB | 2.0 GB | 7.3 GB | 12 GB | Very high |
| Q4_K_M * | 4.68 GB | 2.0 GB | 6.7 GB | 8 GB | High — the default pick |
| Q4_K_L | 5.11 GB | 2.0 GB | 7.1 GB | 8 GB | High |
| Q4_1 | 4.89 GB | 2.0 GB | 6.9 GB | 8 GB | High |
| Q4_K_S | 4.47 GB | 2.0 GB | 6.5 GB | 8 GB | High |
| IQ4_NL | 4.46 GB | 2.0 GB | 6.5 GB | 8 GB | High |
| Q4_0 | 4.46 GB | 2.0 GB | 6.5 GB | 8 GB | High |
| IQ4_XS | 4.25 GB | 2.0 GB | 6.3 GB | 8 GB | High |
| Q3_K_XL | 4.63 GB | 2.0 GB | 6.6 GB | 8 GB | Acceptable — visible loss |
| Q3_K_L | 4.13 GB | 2.0 GB | 6.1 GB | 8 GB | Acceptable — visible loss |
| Q3_K_M | 3.84 GB | 2.0 GB | 5.8 GB | 8 GB | Acceptable — visible loss |
| IQ3_M | 3.63 GB | 2.0 GB | 5.6 GB | 8 GB | Acceptable — visible loss |
| Q3_K_S | 3.51 GB | 2.0 GB | 5.5 GB | 8 GB | Acceptable — visible loss |
| IQ3_XS | 3.38 GB | 2.0 GB | 5.4 GB | 8 GB | Acceptable — visible loss |
| IQ3_XXS | 3.14 GB | 2.0 GB | 5.1 GB | 8 GB | Acceptable — visible loss |
| Q2_K_L | 3.62 GB | 2.0 GB | 5.6 GB | 8 GB | Experimental — not ranked — never recommended |
| Q2_K | 3.06 GB | 2.0 GB | 5.1 GB | 8 GB | Experimental — not ranked — never recommended |
| IQ2_M | 2.84 GB | 2.0 GB | 4.8 GB | 8 GB | Experimental — not ranked — never recommended |
* default pick. Weights = real GGUF file sizes from bartowski/Qwen_Qwen3-8B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.
Best Qwen3 8B quant by memory
| Memory | Recommended quant | Total (16k ctx) |
|---|---|---|
| 8 GB | Q4_K_M | 6.7 GB |
| 12 GB | Q8_0 | 10.1 GB |
| 24 GB | BF16 | 17.3 GB |
Why we don't rank Qwen3 8B's 2-bit quants
Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.
Run Qwen3 8B on your GPU
Frequently asked questions
What is the best quantization of Qwen3 8B?
Q4_K_M is the default pick: 4.68 GB of weights, high — the default pick quality, fitting comfortably in 8 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.
How much memory does Qwen3 8B need?
At Q4_K_M, Qwen3 8B needs 4.68 GB for the weights plus ~2.0 GB of KV-cache at 16k context — about 6.7 GB total, so a 8 GB card or Mac (90% usable budget) runs it comfortably.
Should I use a Q2_K or IQ2 quant of Qwen3 8B?
No. Qwen3 8B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.
Cite this page
ModelFit: Qwen3 8B quantization comparison (real GGUF file sizes). https://modelfit.io/quant-compare/qwen3-8b/ (data probed 2026-09-02, CC BY 4.0).