Qwen3.8 27B quants compared

25 real GGUF builds of Qwen3.8 27B measured from unsloth/Qwen3.8-27B-GGUF — download the Q4_K_M (15.33 GB) unless you know why you need heavier. Probed 2026-09-02.

Download Qwen3.8 27B Q4_K_M (15.33 GB) — it fits 24 GB of memory with 16k context. With Ollama: ollama run qwen3.8:27b

Every Qwen3.8 27B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1650.9 GB1.0 GB51.9 GB64 GBFull precision (lossless)
Q8_K_XL29.3 GB1.0 GB30.3 GB48 GBNear-lossless
Q8_027.05 GB1.0 GB28.1 GB32 GBNear-lossless
Q8_K_L26.12 GB1.0 GB27.1 GB32 GBNear-lossless
Q6_K_XL23.56 GB1.0 GB24.6 GB32 GBExcellent
Q6_K_L22.53 GB1.0 GB23.5 GB32 GBExcellent
Q6_K_M21.5 GB1.0 GB22.5 GB32 GBExcellent
Q6_K20.47 GB1.0 GB21.5 GB24 GBExcellent
Q5_K_XL19.44 GB1.0 GB20.4 GB24 GBVery high
Q5_K_M18.41 GB1.0 GB19.4 GB24 GBVery high
Q5_K_S17.38 GB1.0 GB18.4 GB24 GBVery high
Q4_K_M *15.33 GB1.0 GB16.3 GB24 GBHigh — the default pick
Q4_K_XL16.35 GB1.0 GB17.4 GB24 GBHigh
Q4_116.34 GB1.0 GB17.3 GB24 GBHigh
Q4_016.23 GB1.0 GB17.2 GB24 GBHigh
Q4_K_S14.3 GB1.0 GB15.3 GB24 GBHigh
IQ4_XS13.27 GB1.0 GB14.3 GB16 GBHigh
Q3_K_XL12.24 GB1.0 GB13.2 GB16 GBAcceptable — visible loss
IQ3_S11.21 GB1.0 GB12.2 GB16 GBAcceptable — visible loss
IQ3_XXS10.18 GB1.0 GB11.2 GB16 GBAcceptable — visible loss
Q2_K_XL9.15 GB1.0 GB10.2 GB12 GBExperimental — not ranked — never recommended
IQ2_S7.8 GB1.0 GB8.8 GB12 GBExperimental — not ranked — never recommended
IQ2_XXS6.77 GB1.0 GB7.8 GB12 GBExperimental — not ranked — never recommended
IQ1_M6.27 GB1.0 GB7.3 GB12 GBExperimental — not ranked — never recommended
IQ1_S5.77 GB1.0 GB6.8 GB8 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from unsloth/Qwen3.8-27B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3.8 27B quant by memory

MemoryRecommended quantTotal (16k ctx)
16 GBIQ4_XS14.3 GB
24 GBQ6_K21.5 GB
32 GBQ8_028.1 GB
64 GBBF1651.9 GB

Why we don't rank Qwen3.8 27B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Run Qwen3.8 27B on your GPU

Looking at hardware first? Best hardware for Qwen3.8 27B — cheapest card, best value per dollar, fastest machine, and Mac fit.

Frequently asked questions

What is the best quantization of Qwen3.8 27B?

Q4_K_M is the default pick: 15.33 GB of weights, high — the default pick quality, fitting comfortably in 24 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3.8 27B need?

At Q4_K_M, Qwen3.8 27B needs 15.33 GB for the weights plus ~1.0 GB of KV-cache at 16k context — about 16.3 GB total, so a 24 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3.8 27B?

No. Qwen3.8 27B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3.8 27B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3.8-27b/ (data probed 2026-09-02, CC BY 4.0).