Gemma 4 31B
Gemma 4 31B loads in 20 GB at Q4_K_M and is comfortable from 36 GB of memory, Macs included from the MacBook Pro M5 Pro 48GB. Below: the full memory math, the cheapest card that runs it well, and the fastest machine we track.
PARAMETERS
31B
FORMAT
Q4_K_M
MIN MEMORY
32 GB
BEST FOR
Quality, Coding, Multimodal
Memory math
Weights (Q4_K_M)
20 GB
+ KV cache (16k)
~4.0 GB
Total at 16k
~24.0 GB
Comfortable from
36 GB
KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).
Hardware snapshot
Go deeper
Every Gemma 4 31B quant by real GGUF file sizeAlso tracked: Gemma 4 31B (Q8) (Q8_0, 30.9 GB, min 48 GB)
More Gemma models
Frequently asked questions
How much memory does Gemma 4 31B need?
20 GB for the Q4_K_M weights, plus ~4.0 GB of KV-cache at 16k context — about 24.0 GB total. Comfortable from 36 GB of VRAM or unified memory.
Does Gemma 4 31B run on a Mac?
Yes — from the MacBook Pro M5 Pro 48GB (~13 tok/s est.). Unified memory means the RAM budget is the only limit.
What is the cheapest GPU for Gemma 4 31B?
The AMD Radeon RX 7900 XTX is the cheapest tracked card that runs Gemma 4 31B comfortably — ~28 tok/s est. at ~$700 used.
Cite this page
ModelFit: Gemma 4 31B — specs, memory math and hardware verdicts. https://modelfit.io/models/gemma4-31b/ (dataset updated 2026-09-03, CC BY 4.0).