By Peter · ModelFit · 2026-09-12

Mistral Small VRAM Requirements (2026)

A graphics card box panel printed with the colorful Mistral AI bar logo, representing a 22B model that fits 24GB cards

Mistral Small 22B is the quiet value pick of the mid-size class. At 17 GB loaded, it sits exactly at the boundary between 16GB and 24GB hardware, which is why it is so often mis-sized. It does not fit a 16GB GPU, and it wastes a 32GB one. It belongs on 20GB to 24GB, where it delivers dense-22B coding quality that the smaller MoE models do not quite match. We tested every fit verdict in this guide with the ModelFit engine on September 3, 2026.

TL;DR: Mistral Small 22B loads in ~17 GB at Q4_K_M. Plan for 20GB minimum, 24GB comfortable. It runs at ~44 tok/s est. on an RTX 4090 and ~37 on an RX 7900 XTX. It does not fit 16GB GPUs. On Macs it wants 24GB unified memory minimum.

How much VRAM does Mistral Small 22B need?

QuantApprox. sizeFits 24GB?Fits 16GB?
Q4_K_M~17 GBYes, comfortableNo
Q8_0~24 GBTightNo
Q2/Q3<14 GBYesMarginal, not recommended

Context adds on top, and this is where the dense architecture costs you. As a dense 22B from Mistral AI, Mistral Small carries a standard GQA KV cache of roughly 192 KB per token at fp16. A 32K context alone costs about 6 GB on top of the 17 GB of weights. On a 24GB card, that said, keep contexts at 16K-32K. On 20GB, stay at 8K-16K. Memory you reserve for context is memory the weights can no longer use. The builds we track match the official ones on Mistral's Hugging Face org. You can line the quants up side by side on our quant comparison tool.

Which hardware runs it well?

ModelFit engine estimates at Q4_K_M, labeled est.

HardwareMemoryEst. tok/sVerdict
RTX 4090 24GB24 GB~44 est.Excellent
RX 7900 XTX 24GB24 GB~38 est.Excellent
RTX 3090 24GB24 GB~37 est.Excellent
RX 7900 XT 20GB20 GB~31 est.Comfortable
RTX 5080 16GB16 GB~8 est.Offload, avoid
RTX 5070 12GB12 GB~3 est.Does not fit
Mac Mini M6 24GB24 GB unified~19 est.Comfortable
Mac Mini M6 32GB32 GB unified~19 est.Comfortable with headroom

The 16GB rows are the trap. Mistral Small will load on an RTX 5080 with part of the weights offloaded, at about 8 tok/s est. In practice that is too slow for a model whose entire appeal is interactive coding. For example, a multi-file refactor that should stream out in seconds turns into minutes of watching tokens trickle. That is exactly the experience that makes people conclude local models are not ready. The real problem is a bad fit. Full per-GPU verdicts live on the Mistral Small 22B model page. The same math is reproducible for your own card on our hardware fit checker.

Where does Mistral Small win?

Its niche is dense-model coding quality at the top of the single-GPU range. The tradeoff is worth understanding before you buy hardware around it. The 20B-class MoE models (GPT-OSS 20B, Gemma 4 26B-A4B) are faster at the same memory. They only read their active experts per token. Mistral Small reads all 22B of its weights on every single token it generates. What those dense weights buy in return is consistency. On long multi-file coding sessions, a sparse model occasionally routes a tricky prompt to the wrong experts. Mistral Small's instruction following holds steadier from the first token to the last. If you want the fastest model at 17 GB, pick the MoE. If you want the most predictable dense coder that still fits one card, Mistral Small holds its ground. We tested it against both MoE rivals in the engine, and the speed gap is real — roughly double on the same card.

On 24GB Macs it is one of the better dense picks, at ~19 tok/s est. on an M6 Mac Mini. It is also the largest dense model that tier runs with real context headroom. The same bandwidth ceiling that caps the GPU cards applies here, so the verdict travels across platforms. Comfortable for interactive coding, not built for batch jobs that run unattended all night.

How should you choose between the 20GB and 24GB cards?

The table above compresses a real decision. Mistral Small is one of the few models where 20GB and 24GB are genuinely different experiences, not a rounding error. At 20GB, such as the RX 7900 XT, the model fits with modest context. You live in the 8K-16K range. That covers chat and single-file work, but starts to pinch on repository-scale coding sessions. At 24GB the context budget roughly doubles. That headroom, not the raw speed difference between the cards, is what you are actually paying for.

If you are buying hardware specifically for this model, the used RTX 3090 24GB remains the value answer. Anyone shopping new should price the jump to 32GB before settling. The next model up the ladder wants that tier.

What should you run instead?

FAQ

Can Mistral Small run on 16GB VRAM?

Not well. At 17 GB it exceeds 16GB before context, so it offloads and drops to roughly 8 tok/s est. on an RTX 5080. Use a 12B model on 16GB and step up to Mistral Small at 20GB+.

Is Mistral Small better than GPT-OSS 20B?

Different strengths. GPT-OSS 20B is faster and lighter (about 14 GB vs 17 GB) with stronger general reasoning. Mistral Small is denser and more consistent on long coding sessions. For most users the MoE is the better default; Mistral Small is the dense-model preference.

What is the cheapest card that runs it well?

A used RTX 3090 24GB. It runs Mistral Small at ~37 tok/s est., within 15% of a 4090, for a fraction of the price.

Does it run on a 16GB Mac?

No. The model is 17 GB before macOS takes its share. A 16GB Mac tops out at 12B models. See the 16GB Mac guide.

___

Fit verdicts and speeds in this guide come from the ModelFit engine, re-run on September 3, 2026, with estimates labeled est. How we work: about ModelFit. Spotted a stale figure? Contact us.
What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

The weekly local-AI refresh

New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.

By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.

Have questions? Reach out on X/Twitter