Qwen3 8B quants compared

24 real GGUF builds of Qwen3 8B measured from bartowski/Qwen_Qwen3-8B-GGUF — download the Q4_K_M (4.68 GB) unless you know why you need heavier. Probed 2026-09-02.

Download Qwen3 8B Q4_K_M (4.68 GB) — it fits 8 GB of memory with 16k context. With Ollama: ollama run qwen3:8b-q4_K_M

Every Qwen3 8B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1615.26 GB2.0 GB17.3 GB24 GBFull precision (lossless)
Q8_08.11 GB2.0 GB10.1 GB12 GBNear-lossless
Q6_K_L6.54 GB2.0 GB8.5 GB12 GBExcellent
Q6_K6.26 GB2.0 GB8.3 GB12 GBExcellent
Q5_K_L5.81 GB2.0 GB7.8 GB12 GBVery high
Q5_K_M5.45 GB2.0 GB7.5 GB12 GBVery high
Q5_K_S5.33 GB2.0 GB7.3 GB12 GBVery high
Q4_K_M *4.68 GB2.0 GB6.7 GB8 GBHigh — the default pick
Q4_K_L5.11 GB2.0 GB7.1 GB8 GBHigh
Q4_14.89 GB2.0 GB6.9 GB8 GBHigh
Q4_K_S4.47 GB2.0 GB6.5 GB8 GBHigh
IQ4_NL4.46 GB2.0 GB6.5 GB8 GBHigh
Q4_04.46 GB2.0 GB6.5 GB8 GBHigh
IQ4_XS4.25 GB2.0 GB6.3 GB8 GBHigh
Q3_K_XL4.63 GB2.0 GB6.6 GB8 GBAcceptable — visible loss
Q3_K_L4.13 GB2.0 GB6.1 GB8 GBAcceptable — visible loss
Q3_K_M3.84 GB2.0 GB5.8 GB8 GBAcceptable — visible loss
IQ3_M3.63 GB2.0 GB5.6 GB8 GBAcceptable — visible loss
Q3_K_S3.51 GB2.0 GB5.5 GB8 GBAcceptable — visible loss
IQ3_XS3.38 GB2.0 GB5.4 GB8 GBAcceptable — visible loss
IQ3_XXS3.14 GB2.0 GB5.1 GB8 GBAcceptable — visible loss
Q2_K_L3.62 GB2.0 GB5.6 GB8 GBExperimental — not ranked — never recommended
Q2_K3.06 GB2.0 GB5.1 GB8 GBExperimental — not ranked — never recommended
IQ2_M2.84 GB2.0 GB4.8 GB8 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from bartowski/Qwen_Qwen3-8B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3 8B quant by memory

MemoryRecommended quantTotal (16k ctx)
8 GBQ4_K_M6.7 GB
12 GBQ8_010.1 GB
24 GBBF1617.3 GB

Why we don't rank Qwen3 8B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Run Qwen3 8B on your GPU

Looking at hardware first? Best hardware for Qwen3 8B — cheapest card, best value per dollar, fastest machine, and Mac fit.

Frequently asked questions

What is the best quantization of Qwen3 8B?

Q4_K_M is the default pick: 4.68 GB of weights, high — the default pick quality, fitting comfortably in 8 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3 8B need?

At Q4_K_M, Qwen3 8B needs 4.68 GB for the weights plus ~2.0 GB of KV-cache at 16k context — about 6.7 GB total, so a 8 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3 8B?

No. Qwen3 8B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3 8B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3-8b/ (data probed 2026-09-02, CC BY 4.0).