Gemma 3 4B Instruct quants compared
26 GGUF builds by real file size, probed from unsloth/gemma-3-4b-it-GGUF on Hugging Face (2026-09-02). 4B params.
Download Gemma 3 4B Instruct Q4_K_M (2.32 GB) — it fits 8 GB of memory with 16k context. With Ollama: ollama run gemma3:4b
Every Gemma 3 4B Instruct quant by real file size
| Quant | Weights | + KV (16k) | Total | Fits comfortably in | Quality |
|---|---|---|---|---|---|
| BF16 | 7.23 GB | 1.8 GB | 9.0 GB | 12 GB | Full precision (lossless) |
| Q8_K_XL | 4.81 GB | 1.8 GB | 6.6 GB | 8 GB | Near-lossless |
| Q8_0 | 3.85 GB | 1.8 GB | 5.6 GB | 8 GB | Near-lossless |
| Q6_K_XL | 3.32 GB | 1.8 GB | 5.1 GB | 8 GB | Excellent |
| Q6_K | 2.97 GB | 1.8 GB | 4.7 GB | 8 GB | Excellent |
| Q5_K_M | 2.64 GB | 1.8 GB | 4.4 GB | 8 GB | Very high |
| Q5_K_XL | 2.64 GB | 1.8 GB | 4.4 GB | 8 GB | Very high |
| Q5_K_S | 2.57 GB | 1.8 GB | 4.3 GB | 8 GB | Very high |
| Q4_K_M * | 2.32 GB | 1.8 GB | 4.1 GB | 8 GB | High — the default pick |
| Q4_1 | 2.39 GB | 1.8 GB | 4.1 GB | 8 GB | High |
| Q4_K_XL | 2.37 GB | 1.8 GB | 4.1 GB | 8 GB | High |
| Q4_0 | 2.21 GB | 1.8 GB | 4.0 GB | 8 GB | High |
| Q4_K_S | 2.21 GB | 1.8 GB | 4.0 GB | 8 GB | High |
| IQ4_NL | 2.2 GB | 1.8 GB | 4.0 GB | 8 GB | High |
| IQ4_XS | 2.11 GB | 1.8 GB | 3.9 GB | 8 GB | High |
| Q3_K_XL | 2 GB | 1.8 GB | 3.8 GB | 8 GB | Acceptable — visible loss |
| Q3_K_M | 1.95 GB | 1.8 GB | 3.7 GB | 8 GB | Acceptable — visible loss |
| Q3_K_S | 1.8 GB | 1.8 GB | 3.5 GB | 8 GB | Acceptable — visible loss |
| IQ3_XXS | 1.59 GB | 1.8 GB | 3.3 GB | 8 GB | Acceptable — visible loss |
| Q2_K_XL | 1.65 GB | 1.8 GB | 3.4 GB | 8 GB | Experimental — not ranked — never recommended |
| Q2_K | 1.61 GB | 1.8 GB | 3.4 GB | 8 GB | Experimental — not ranked — never recommended |
| Q2_K_L | 1.61 GB | 1.8 GB | 3.4 GB | 8 GB | Experimental — not ranked — never recommended |
| IQ2_M | 1.46 GB | 1.8 GB | 3.2 GB | 8 GB | Experimental — not ranked — never recommended |
| IQ2_XXS | 1.25 GB | 1.8 GB | 3.0 GB | 8 GB | Experimental — not ranked — never recommended |
| IQ1_M | 1.15 GB | 1.8 GB | 2.9 GB | 8 GB | Experimental — not ranked — never recommended |
| IQ1_S | 1.1 GB | 1.8 GB | 2.9 GB | 8 GB | Experimental — not ranked — never recommended |
* default pick. Weights = real GGUF file sizes from unsloth/gemma-3-4b-it-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.
Best Gemma 3 4B Instruct quant by memory
| Memory | Recommended quant | Total (16k ctx) |
|---|---|---|
| 8 GB | Q8_0 | 5.6 GB |
| 12 GB | BF16 | 9.0 GB |
Why we don't rank Gemma 3 4B Instruct's 2-bit quants
Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.
Frequently asked questions
What is the best quantization of Gemma 3 4B Instruct?
Q4_K_M is the default pick: 2.32 GB of weights, high — the default pick quality, fitting comfortably in 8 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.
How much memory does Gemma 3 4B Instruct need?
At Q4_K_M, Gemma 3 4B Instruct needs 2.32 GB for the weights plus ~1.8 GB of KV-cache at 16k context — about 4.1 GB total, so a 8 GB card or Mac (90% usable budget) runs it comfortably.
Should I use a Q2_K or IQ2 quant of Gemma 3 4B Instruct?
No. Gemma 3 4B Instruct at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.
Cite this page
ModelFit: Gemma 3 4B Instruct quantization comparison (real GGUF file sizes). https://modelfit.io/quant-compare/gemma3-4b/ (data probed 2026-09-02, CC BY 4.0).