Llama 4 Scout quants compared
27 GGUF builds by real file size, probed from unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF on Hugging Face (2026-09-02). 109B params, 17B active.
Download Llama 4 Scout Q4_K_M (60.87 GB across 2 shards) — it fits 96 GB of memory with 16k context. With Ollama: ollama run llama4:scout
Every Llama 4 Scout quant by real file size
| Quant | Weights | + KV (16k) | Total | Fits comfortably in | Quality |
|---|---|---|---|---|---|
| BF16 | 200.76 GB | 6.0 GB | 206.8 GB | >128 GB | Full precision (lossless) |
| Q8_K_XL | 119.38 GB | 6.0 GB | 125.4 GB | >128 GB | Near-lossless |
| Q8_0 | 106.67 GB | 6.0 GB | 112.7 GB | 128 GB | Near-lossless |
| Q6_K_XL | 87.61 GB | 6.0 GB | 93.6 GB | 128 GB | Excellent |
| Q6_K | 82.36 GB | 6.0 GB | 88.4 GB | 128 GB | Excellent |
| Q5_K_XL | 73.71 GB | 6.0 GB | 79.7 GB | 96 GB | Very high |
| Q5_K_M | 71.29 GB | 6.0 GB | 77.3 GB | 96 GB | Very high |
| Q5_K_S | 69.16 GB | 6.0 GB | 75.2 GB | 96 GB | Very high |
| Q4_K_M * | 60.87 GB | 6.0 GB | 66.9 GB | 96 GB | High — the default pick |
| Q4_1 | 62.94 GB | 6.0 GB | 68.9 GB | 96 GB | High |
| Q4_K_XL | 57.74 GB | 6.0 GB | 63.7 GB | 96 GB | High |
| Q4_K_S | 57.23 GB | 6.0 GB | 63.2 GB | 96 GB | High |
| Q4_0 | 56.98 GB | 6.0 GB | 63.0 GB | 96 GB | High |
| IQ4_NL | 56.76 GB | 6.0 GB | 62.8 GB | 96 GB | High |
| IQ4_XS | 53.69 GB | 6.0 GB | 59.7 GB | 96 GB | High |
| Q3_K_M | 48.2 GB | 6.0 GB | 54.2 GB | 64 GB | Acceptable — visible loss |
| Q3_K_XL | 45.65 GB | 6.0 GB | 51.6 GB | 64 GB | Acceptable — visible loss |
| Q3_K_S | 43.53 GB | 6.0 GB | 49.5 GB | 64 GB | Acceptable — visible loss |
| IQ3_XXS | 42.59 GB | 6.0 GB | 48.6 GB | 64 GB | Acceptable — visible loss |
| Q2_K_XL | 39.47 GB | 6.0 GB | 45.5 GB | 64 GB | Experimental — not ranked — never recommended |
| Q2_K_L | 37.07 GB | 6.0 GB | 43.1 GB | 48 GB | Experimental — not ranked — never recommended |
| Q2_K | 36.85 GB | 6.0 GB | 42.9 GB | 48 GB | Experimental — not ranked — never recommended |
| IQ2_M | 36.39 GB | 6.0 GB | 42.4 GB | 48 GB | Experimental — not ranked — never recommended |
| IQ2_XXS | 34.83 GB | 6.0 GB | 40.8 GB | 48 GB | Experimental — not ranked — never recommended |
| IQ1_M | 32.59 GB | 6.0 GB | 38.6 GB | 48 GB | Experimental — not ranked — never recommended |
| IQ1_S | 30.24 GB | 6.0 GB | 36.2 GB | 48 GB | Experimental — not ranked — never recommended |
| TQ1_0 | 27.25 GB | 6.0 GB | 33.3 GB | 48 GB | Experimental — not ranked — never recommended |
* default pick. Weights = real GGUF file sizes from unsloth/Llama-4-Scout-17B-16E-Instruct-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.
Best Llama 4 Scout quant by memory
| Memory | Recommended quant | Total (16k ctx) |
|---|---|---|
| 64 GB | Q3_K_M | 54.2 GB |
| 96 GB | Q5_K_M | 77.3 GB |
| 128 GB | Q8_0 | 112.7 GB |
Why we don't rank Llama 4 Scout's 2-bit quants
Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.
Frequently asked questions
What is the best quantization of Llama 4 Scout?
Q4_K_M is the default pick: 60.87 GB of weights, high — the default pick quality, fitting comfortably in 96 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.
How much memory does Llama 4 Scout need?
At Q4_K_M, Llama 4 Scout needs 60.87 GB for the weights plus ~6.0 GB of KV-cache at 16k context — about 66.9 GB total, so a 96 GB card or Mac (90% usable budget) runs it comfortably.
Should I use a Q2_K or IQ2 quant of Llama 4 Scout?
No. Llama 4 Scout at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.
Cite this page
ModelFit: Llama 4 Scout quantization comparison (real GGUF file sizes). https://modelfit.io/quant-compare/llama4-scout/ (data probed 2026-09-02, CC BY 4.0).