Qwen3-Next 80B-A3B quants compared

26 GGUF builds by real file size, probed from unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF on Hugging Face (2026-09-02). 80B params, 3B active.

Download Qwen3-Next 80B-A3B Q4_K_M (45.17 GB) — it fits 64 GB of memory with 16k context. With Ollama: ollama run qwen3-next:80b

Every Qwen3-Next 80B-A3B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF16148.51 GB0.4 GB148.9 GB>128 GBFull precision (lossless)
Q8_K_XL86.68 GB0.4 GB87.1 GB128 GBNear-lossless
Q8_078.99 GB0.4 GB79.4 GB96 GBNear-lossless
Q6_K_XL63.81 GB0.4 GB64.2 GB96 GBExcellent
Q6_K61.04 GB0.4 GB61.4 GB96 GBExcellent
Q5_K_M52.91 GB0.4 GB53.3 GB64 GBVery high
Q5_K_XL52.77 GB0.4 GB53.1 GB64 GBVery high
Q5_K_S51.24 GB0.4 GB51.6 GB64 GBVery high
Q4_K_M *45.17 GB0.4 GB45.5 GB64 GBHigh — the default pick
Q4_146.62 GB0.4 GB47.0 GB64 GBHigh
Q4_K_XL42.9 GB0.4 GB43.3 GB64 GBHigh
Q4_K_S42.38 GB0.4 GB42.8 GB48 GBHigh
Q4_042.2 GB0.4 GB42.6 GB48 GBHigh
IQ4_NL42.01 GB0.4 GB42.4 GB48 GBHigh
IQ4_XS39.72 GB0.4 GB40.1 GB48 GBHigh
Q3_K_M35.67 GB0.4 GB36.0 GB48 GBAcceptable — visible loss
Q3_K_XL33.19 GB0.4 GB33.6 GB48 GBAcceptable — visible loss
Q3_K_S32.21 GB0.4 GB32.6 GB48 GBAcceptable — visible loss
IQ3_XXS30.82 GB0.4 GB31.2 GB48 GBAcceptable — visible loss
Q2_K_XL28.06 GB0.4 GB28.4 GB32 GBExperimental — not ranked — never recommended
Q2_K_L27.24 GB0.4 GB27.6 GB32 GBExperimental — not ranked — never recommended
Q2_K27.17 GB0.4 GB27.5 GB32 GBExperimental — not ranked — never recommended
IQ2_XXS24.41 GB0.4 GB24.8 GB32 GBExperimental — not ranked — never recommended
IQ1_M22.73 GB0.4 GB23.1 GB32 GBExperimental — not ranked — never recommended
IQ1_S21.33 GB0.4 GB21.7 GB32 GBExperimental — not ranked — never recommended
TQ1_019.06 GB0.4 GB19.4 GB24 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from unsloth/Qwen3-Next-80B-A3B-Instruct-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3-Next 80B-A3B quant by memory

MemoryRecommended quantTotal (16k ctx)
48 GBQ4_042.6 GB
64 GBQ5_K_M53.3 GB
96 GBQ8_079.4 GB

Why we don't rank Qwen3-Next 80B-A3B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of Qwen3-Next 80B-A3B?

Q4_K_M is the default pick: 45.17 GB of weights, high — the default pick quality, fitting comfortably in 64 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3-Next 80B-A3B need?

At Q4_K_M, Qwen3-Next 80B-A3B needs 45.17 GB for the weights plus ~0.4 GB of KV-cache at 16k context — about 45.5 GB total, so a 64 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3-Next 80B-A3B?

No. Qwen3-Next 80B-A3B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3-Next 80B-A3B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3-next-80b-a3b/ (data probed 2026-09-02, CC BY 4.0).