Llama 4 Scout quants compared

27 GGUF builds by real file size, probed from unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF on Hugging Face (2026-09-02). 109B params, 17B active.

Download Llama 4 Scout Q4_K_M (60.87 GB across 2 shards) — it fits 96 GB of memory with 16k context. With Ollama: ollama run llama4:scout

Every Llama 4 Scout quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF16200.76 GB6.0 GB206.8 GB>128 GBFull precision (lossless)
Q8_K_XL119.38 GB6.0 GB125.4 GB>128 GBNear-lossless
Q8_0106.67 GB6.0 GB112.7 GB128 GBNear-lossless
Q6_K_XL87.61 GB6.0 GB93.6 GB128 GBExcellent
Q6_K82.36 GB6.0 GB88.4 GB128 GBExcellent
Q5_K_XL73.71 GB6.0 GB79.7 GB96 GBVery high
Q5_K_M71.29 GB6.0 GB77.3 GB96 GBVery high
Q5_K_S69.16 GB6.0 GB75.2 GB96 GBVery high
Q4_K_M *60.87 GB6.0 GB66.9 GB96 GBHigh — the default pick
Q4_162.94 GB6.0 GB68.9 GB96 GBHigh
Q4_K_XL57.74 GB6.0 GB63.7 GB96 GBHigh
Q4_K_S57.23 GB6.0 GB63.2 GB96 GBHigh
Q4_056.98 GB6.0 GB63.0 GB96 GBHigh
IQ4_NL56.76 GB6.0 GB62.8 GB96 GBHigh
IQ4_XS53.69 GB6.0 GB59.7 GB96 GBHigh
Q3_K_M48.2 GB6.0 GB54.2 GB64 GBAcceptable — visible loss
Q3_K_XL45.65 GB6.0 GB51.6 GB64 GBAcceptable — visible loss
Q3_K_S43.53 GB6.0 GB49.5 GB64 GBAcceptable — visible loss
IQ3_XXS42.59 GB6.0 GB48.6 GB64 GBAcceptable — visible loss
Q2_K_XL39.47 GB6.0 GB45.5 GB64 GBExperimental — not ranked — never recommended
Q2_K_L37.07 GB6.0 GB43.1 GB48 GBExperimental — not ranked — never recommended
Q2_K36.85 GB6.0 GB42.9 GB48 GBExperimental — not ranked — never recommended
IQ2_M36.39 GB6.0 GB42.4 GB48 GBExperimental — not ranked — never recommended
IQ2_XXS34.83 GB6.0 GB40.8 GB48 GBExperimental — not ranked — never recommended
IQ1_M32.59 GB6.0 GB38.6 GB48 GBExperimental — not ranked — never recommended
IQ1_S30.24 GB6.0 GB36.2 GB48 GBExperimental — not ranked — never recommended
TQ1_027.25 GB6.0 GB33.3 GB48 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Llama 4 Scout quant by memory

MemoryRecommended quantTotal (16k ctx)
64 GBQ3_K_M54.2 GB
96 GBQ5_K_M77.3 GB
128 GBQ8_0112.7 GB

Why we don't rank Llama 4 Scout's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of Llama 4 Scout?

Q4_K_M is the default pick: 60.87 GB of weights, high — the default pick quality, fitting comfortably in 96 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Llama 4 Scout need?

At Q4_K_M, Llama 4 Scout needs 60.87 GB for the weights plus ~6.0 GB of KV-cache at 16k context — about 66.9 GB total, so a 96 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Llama 4 Scout?

No. Llama 4 Scout at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Llama 4 Scout quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/llama4-scout/ (data probed 2026-09-02, CC BY 4.0).