Qwen3.8 27B quants compared
25 GGUF builds by real file size, probed from unsloth/Qwen3.8-27B-GGUF on Hugging Face (2026-09-02). 27B params.
Download Qwen3.8 27B Q4_K_M (15.33 GB) — it fits 24 GB of memory with 16k context. With Ollama: ollama run qwen3.8:27b
Every Qwen3.8 27B quant by real file size
| Quant | Weights | + KV (16k) | Total | Fits comfortably in | Quality |
|---|---|---|---|---|---|
| BF16 | 50.9 GB | 1.0 GB | 51.9 GB | 64 GB | Full precision (lossless) |
| Q8_K_XL | 29.3 GB | 1.0 GB | 30.3 GB | 48 GB | Near-lossless |
| Q8_0 | 27.05 GB | 1.0 GB | 28.1 GB | 32 GB | Near-lossless |
| Q8_K_L | 26.12 GB | 1.0 GB | 27.1 GB | 32 GB | Near-lossless |
| Q6_K_XL | 23.56 GB | 1.0 GB | 24.6 GB | 32 GB | Excellent |
| Q6_K_L | 22.53 GB | 1.0 GB | 23.5 GB | 32 GB | Excellent |
| Q6_K_M | 21.5 GB | 1.0 GB | 22.5 GB | 32 GB | Excellent |
| Q6_K | 20.47 GB | 1.0 GB | 21.5 GB | 24 GB | Excellent |
| Q5_K_XL | 19.44 GB | 1.0 GB | 20.4 GB | 24 GB | Very high |
| Q5_K_M | 18.41 GB | 1.0 GB | 19.4 GB | 24 GB | Very high |
| Q5_K_S | 17.38 GB | 1.0 GB | 18.4 GB | 24 GB | Very high |
| Q4_K_M * | 15.33 GB | 1.0 GB | 16.3 GB | 24 GB | High — the default pick |
| Q4_K_XL | 16.35 GB | 1.0 GB | 17.4 GB | 24 GB | High |
| Q4_1 | 16.34 GB | 1.0 GB | 17.3 GB | 24 GB | High |
| Q4_0 | 16.23 GB | 1.0 GB | 17.2 GB | 24 GB | High |
| Q4_K_S | 14.3 GB | 1.0 GB | 15.3 GB | 24 GB | High |
| IQ4_XS | 13.27 GB | 1.0 GB | 14.3 GB | 16 GB | High |
| Q3_K_XL | 12.24 GB | 1.0 GB | 13.2 GB | 16 GB | Acceptable — visible loss |
| IQ3_S | 11.21 GB | 1.0 GB | 12.2 GB | 16 GB | Acceptable — visible loss |
| IQ3_XXS | 10.18 GB | 1.0 GB | 11.2 GB | 16 GB | Acceptable — visible loss |
| Q2_K_XL | 9.15 GB | 1.0 GB | 10.2 GB | 12 GB | Experimental — not ranked — never recommended |
| IQ2_S | 7.8 GB | 1.0 GB | 8.8 GB | 12 GB | Experimental — not ranked — never recommended |
| IQ2_XXS | 6.77 GB | 1.0 GB | 7.8 GB | 12 GB | Experimental — not ranked — never recommended |
| IQ1_M | 6.27 GB | 1.0 GB | 7.3 GB | 12 GB | Experimental — not ranked — never recommended |
| IQ1_S | 5.77 GB | 1.0 GB | 6.8 GB | 8 GB | Experimental — not ranked — never recommended |
* default pick. Weights = real GGUF file sizes from unsloth/Qwen3.8-27B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.
Best Qwen3.8 27B quant by memory
| Memory | Recommended quant | Total (16k ctx) |
|---|---|---|
| 16 GB | IQ4_XS | 14.3 GB |
| 24 GB | Q6_K | 21.5 GB |
| 32 GB | Q8_0 | 28.1 GB |
| 64 GB | BF16 | 51.9 GB |
Why we don't rank Qwen3.8 27B's 2-bit quants
Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.
Run Qwen3.8 27B on your GPU
Frequently asked questions
What is the best quantization of Qwen3.8 27B?
Q4_K_M is the default pick: 15.33 GB of weights, high — the default pick quality, fitting comfortably in 24 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.
How much memory does Qwen3.8 27B need?
At Q4_K_M, Qwen3.8 27B needs 15.33 GB for the weights plus ~1.0 GB of KV-cache at 16k context — about 16.3 GB total, so a 24 GB card or Mac (90% usable budget) runs it comfortably.
Should I use a Q2_K or IQ2 quant of Qwen3.8 27B?
No. Qwen3.8 27B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.
Cite this page
ModelFit: Qwen3.8 27B quantization comparison (real GGUF file sizes). https://modelfit.io/quant-compare/qwen3.8-27b/ (data probed 2026-09-02, CC BY 4.0).