Qwen3 8B quants compared

24 GGUF builds by real file size, probed from bartowski/Qwen_Qwen3-8B-GGUF on Hugging Face (2026-09-02). 8B params.

Download Qwen3 8B Q4_K_M (4.68 GB) — it fits 8 GB of memory with 16k context. With Ollama: ollama run qwen3:8b-q4_K_M

Every Qwen3 8B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1615.26 GB2.0 GB17.3 GB24 GBFull precision (lossless)
Q8_08.11 GB2.0 GB10.1 GB12 GBNear-lossless
Q6_K_L6.54 GB2.0 GB8.5 GB12 GBExcellent
Q6_K6.26 GB2.0 GB8.3 GB12 GBExcellent
Q5_K_L5.81 GB2.0 GB7.8 GB12 GBVery high
Q5_K_M5.45 GB2.0 GB7.5 GB12 GBVery high
Q5_K_S5.33 GB2.0 GB7.3 GB12 GBVery high
Q4_K_M *4.68 GB2.0 GB6.7 GB8 GBHigh — the default pick
Q4_K_L5.11 GB2.0 GB7.1 GB8 GBHigh
Q4_14.89 GB2.0 GB6.9 GB8 GBHigh
Q4_K_S4.47 GB2.0 GB6.5 GB8 GBHigh
IQ4_NL4.46 GB2.0 GB6.5 GB8 GBHigh
Q4_04.46 GB2.0 GB6.5 GB8 GBHigh
IQ4_XS4.25 GB2.0 GB6.3 GB8 GBHigh
Q3_K_XL4.63 GB2.0 GB6.6 GB8 GBAcceptable — visible loss
Q3_K_L4.13 GB2.0 GB6.1 GB8 GBAcceptable — visible loss
Q3_K_M3.84 GB2.0 GB5.8 GB8 GBAcceptable — visible loss
IQ3_M3.63 GB2.0 GB5.6 GB8 GBAcceptable — visible loss
Q3_K_S3.51 GB2.0 GB5.5 GB8 GBAcceptable — visible loss
IQ3_XS3.38 GB2.0 GB5.4 GB8 GBAcceptable — visible loss
IQ3_XXS3.14 GB2.0 GB5.1 GB8 GBAcceptable — visible loss
Q2_K_L3.62 GB2.0 GB5.6 GB8 GBExperimental — not ranked — never recommended
Q2_K3.06 GB2.0 GB5.1 GB8 GBExperimental — not ranked — never recommended
IQ2_M2.84 GB2.0 GB4.8 GB8 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from bartowski/Qwen_Qwen3-8B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3 8B quant by memory

MemoryRecommended quantTotal (16k ctx)
8 GBQ4_K_M6.7 GB
12 GBQ8_010.1 GB
24 GBBF1617.3 GB

Why we don't rank Qwen3 8B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Run Qwen3 8B on your GPU

Frequently asked questions

What is the best quantization of Qwen3 8B?

Q4_K_M is the default pick: 4.68 GB of weights, high — the default pick quality, fitting comfortably in 8 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3 8B need?

At Q4_K_M, Qwen3 8B needs 4.68 GB for the weights plus ~2.0 GB of KV-cache at 16k context — about 6.7 GB total, so a 8 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3 8B?

No. Qwen3 8B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3 8B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3-8b/ (data probed 2026-09-02, CC BY 4.0).