Qwen3.8 27B quants compared

25 GGUF builds by real file size, probed from unsloth/Qwen3.8-27B-GGUF on Hugging Face (2026-09-02). 27B params.

Download Qwen3.8 27B Q4_K_M (15.33 GB) — it fits 24 GB of memory with 16k context. With Ollama: ollama run qwen3.8:27b

Every Qwen3.8 27B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1650.9 GB1.0 GB51.9 GB64 GBFull precision (lossless)
Q8_K_XL29.3 GB1.0 GB30.3 GB48 GBNear-lossless
Q8_027.05 GB1.0 GB28.1 GB32 GBNear-lossless
Q8_K_L26.12 GB1.0 GB27.1 GB32 GBNear-lossless
Q6_K_XL23.56 GB1.0 GB24.6 GB32 GBExcellent
Q6_K_L22.53 GB1.0 GB23.5 GB32 GBExcellent
Q6_K_M21.5 GB1.0 GB22.5 GB32 GBExcellent
Q6_K20.47 GB1.0 GB21.5 GB24 GBExcellent
Q5_K_XL19.44 GB1.0 GB20.4 GB24 GBVery high
Q5_K_M18.41 GB1.0 GB19.4 GB24 GBVery high
Q5_K_S17.38 GB1.0 GB18.4 GB24 GBVery high
Q4_K_M *15.33 GB1.0 GB16.3 GB24 GBHigh — the default pick
Q4_K_XL16.35 GB1.0 GB17.4 GB24 GBHigh
Q4_116.34 GB1.0 GB17.3 GB24 GBHigh
Q4_016.23 GB1.0 GB17.2 GB24 GBHigh
Q4_K_S14.3 GB1.0 GB15.3 GB24 GBHigh
IQ4_XS13.27 GB1.0 GB14.3 GB16 GBHigh
Q3_K_XL12.24 GB1.0 GB13.2 GB16 GBAcceptable — visible loss
IQ3_S11.21 GB1.0 GB12.2 GB16 GBAcceptable — visible loss
IQ3_XXS10.18 GB1.0 GB11.2 GB16 GBAcceptable — visible loss
Q2_K_XL9.15 GB1.0 GB10.2 GB12 GBExperimental — not ranked — never recommended
IQ2_S7.8 GB1.0 GB8.8 GB12 GBExperimental — not ranked — never recommended
IQ2_XXS6.77 GB1.0 GB7.8 GB12 GBExperimental — not ranked — never recommended
IQ1_M6.27 GB1.0 GB7.3 GB12 GBExperimental — not ranked — never recommended
IQ1_S5.77 GB1.0 GB6.8 GB8 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from unsloth/Qwen3.8-27B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3.8 27B quant by memory

MemoryRecommended quantTotal (16k ctx)
16 GBIQ4_XS14.3 GB
24 GBQ6_K21.5 GB
32 GBQ8_028.1 GB
64 GBBF1651.9 GB

Why we don't rank Qwen3.8 27B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Run Qwen3.8 27B on your GPU

Frequently asked questions

What is the best quantization of Qwen3.8 27B?

Q4_K_M is the default pick: 15.33 GB of weights, high — the default pick quality, fitting comfortably in 24 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3.8 27B need?

At Q4_K_M, Qwen3.8 27B needs 15.33 GB for the weights plus ~1.0 GB of KV-cache at 16k context — about 16.3 GB total, so a 24 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3.8 27B?

No. Qwen3.8 27B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3.8 27B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3.8-27b/ (data probed 2026-09-02, CC BY 4.0).