Gemma 4 12B quants compared

22 GGUF builds by real file size, probed from unsloth/gemma-4-12b-it-GGUF on Hugging Face (2026-09-02). 12B params.

Download Gemma 4 12B Q4_K_M (6.63 GB) — it fits 12 GB of memory with 16k context. With Ollama: ollama run gemma4:12b

Every Gemma 4 12B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1623 GB3.0 GB26.0 GB32 GBFull precision (lossless)
F160.8 GB3.0 GB3.8 GB8 GBFull precision (lossless)
Q8_K_XL12.7 GB3.0 GB15.7 GB24 GBNear-lossless
Q8_012.23 GB3.0 GB15.2 GB24 GBNear-lossless
Q6_K_XL9.95 GB3.0 GB12.9 GB16 GBExcellent
Q6_K9.11 GB3.0 GB12.1 GB16 GBExcellent
Q5_K_XL8.01 GB3.0 GB11.0 GB16 GBVery high
Q5_K_M7.84 GB3.0 GB10.8 GB16 GBVery high
Q5_K_S7.64 GB3.0 GB10.6 GB12 GBVery high
Q4_K_M *6.63 GB3.0 GB9.6 GB12 GBHigh — the default pick
Q4_16.89 GB3.0 GB9.9 GB12 GBHigh
Q4_K_XL6.86 GB3.0 GB9.9 GB12 GBHigh
Q4_K_S6.3 GB3.0 GB9.3 GB12 GBHigh
Q4_06.28 GB3.0 GB9.3 GB12 GBHigh
IQ4_NL6.26 GB3.0 GB9.3 GB12 GBHigh
IQ4_XS5.94 GB3.0 GB8.9 GB12 GBHigh
Q3_K_XL5.61 GB3.0 GB8.6 GB12 GBAcceptable — visible loss
Q3_K_M5.3 GB3.0 GB8.3 GB12 GBAcceptable — visible loss
Q3_K_S4.78 GB3.0 GB7.8 GB12 GBAcceptable — visible loss
IQ3_XXS4.32 GB3.0 GB7.3 GB12 GBAcceptable — visible loss
Q2_K_XL4.34 GB3.0 GB7.3 GB12 GBExperimental — not ranked — never recommended
IQ2_M3.92 GB3.0 GB6.9 GB8 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from unsloth/gemma-4-12b-it-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Gemma 4 12B quant by memory

MemoryRecommended quantTotal (16k ctx)
8 GBF163.8 GB
32 GBBF1626.0 GB

Why we don't rank Gemma 4 12B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of Gemma 4 12B?

Q4_K_M is the default pick: 6.63 GB of weights, high — the default pick quality, fitting comfortably in 12 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Gemma 4 12B need?

At Q4_K_M, Gemma 4 12B needs 6.63 GB for the weights plus ~3.0 GB of KV-cache at 16k context — about 9.6 GB total, so a 12 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Gemma 4 12B?

No. Gemma 4 12B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Gemma 4 12B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/gemma4-12b/ (data probed 2026-09-02, CC BY 4.0).