Gemma 4 31B quants compared

28 GGUF builds by real file size, probed from bartowski/google_gemma-4-31b-it-GGUF on Hugging Face (2026-09-02). 31B params.

Download Gemma 4 31B Q4_K_M (18.25 GB) — it fits 32 GB of memory with 16k context. With Ollama: ollama run gemma4:31b

Every Gemma 4 31B quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1657.2 GB4.0 GB61.2 GB96 GBFull precision (lossless)
Q8_030.87 GB4.0 GB34.9 GB48 GBNear-lossless
Q6_K_L25.21 GB4.0 GB29.2 GB48 GBExcellent
Q6_K24.89 GB4.0 GB28.9 GB48 GBExcellent
Q5_K_L21.37 GB4.0 GB25.4 GB32 GBVery high
Q5_K_M21.06 GB4.0 GB25.1 GB32 GBVery high
Q5_K_S20.03 GB4.0 GB24.0 GB32 GBVery high
Q4_K_M *18.25 GB4.0 GB22.3 GB32 GBHigh — the default pick
Q4_K_L18.57 GB4.0 GB22.6 GB32 GBHigh
Q4_118.41 GB4.0 GB22.4 GB32 GBHigh
Q4_017.16 GB4.0 GB21.2 GB24 GBHigh
Q4_K_S16.95 GB4.0 GB20.9 GB24 GBHigh
IQ4_NL16.79 GB4.0 GB20.8 GB24 GBHigh
IQ4_XS15.98 GB4.0 GB20.0 GB24 GBHigh
Q3_K_XL15.97 GB4.0 GB20.0 GB24 GBAcceptable — visible loss
Q3_K_L15.66 GB4.0 GB19.7 GB24 GBAcceptable — visible loss
Q3_K_M14.82 GB4.0 GB18.8 GB24 GBAcceptable — visible loss
IQ3_M14.09 GB4.0 GB18.1 GB24 GBAcceptable — visible loss
Q3_K_S13.34 GB4.0 GB17.3 GB24 GBAcceptable — visible loss
IQ3_XS12.89 GB4.0 GB16.9 GB24 GBAcceptable — visible loss
IQ3_XXS12.09 GB4.0 GB16.1 GB24 GBAcceptable — visible loss
Q2_K_L12.08 GB4.0 GB16.1 GB24 GBExperimental — not ranked — never recommended
IQ2_M11.78 GB4.0 GB15.8 GB24 GBExperimental — not ranked — never recommended
Q2_K11.76 GB4.0 GB15.8 GB24 GBExperimental — not ranked — never recommended
IQ2_S11.25 GB4.0 GB15.3 GB24 GBExperimental — not ranked — never recommended
IQ2_XS10.71 GB4.0 GB14.7 GB24 GBExperimental — not ranked — never recommended
IQ2_XXS10.09 GB4.0 GB14.1 GB16 GBExperimental — not ranked — never recommended
IQ1_M9.42 GB4.0 GB13.4 GB16 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from bartowski/google_gemma-4-31b-it-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Gemma 4 31B quant by memory

MemoryRecommended quantTotal (16k ctx)
24 GBQ4_021.2 GB
32 GBQ5_K_M25.1 GB
48 GBQ8_034.9 GB
96 GBBF1661.2 GB

Why we don't rank Gemma 4 31B's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of Gemma 4 31B?

Q4_K_M is the default pick: 18.25 GB of weights, high — the default pick quality, fitting comfortably in 32 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Gemma 4 31B need?

At Q4_K_M, Gemma 4 31B needs 18.25 GB for the weights plus ~4.0 GB of KV-cache at 16k context — about 22.3 GB total, so a 32 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Gemma 4 31B?

No. Gemma 4 31B at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Gemma 4 31B quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/gemma4-31b/ (data probed 2026-09-02, CC BY 4.0).