Mistral Small 3.1 (Q8)

Mistral Small 3.1 (Q8) loads in 23.3 GB at Q8_0 and is comfortable from 48 GB of memory, Macs included from the MacBook Pro M5 Pro 48GB. Below: the full memory math, the cheapest card that runs it well, and the fastest machine we track.

PARAMETERS
24B
FORMAT
Q8_0
MIN MEMORY
36 GB
BEST FOR
Chat, Coding

Memory math

Weights (Q8_0)
23.3 GB
+ KV cache (16k)
~4.0 GB
Total at 16k
~27.3 GB
Comfortable from
48 GB

KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).

Hardware snapshot

Cheapest GPU
~7 tok/s est. — ~$3,847 used (as of 2026-07-31)
Fastest
~35 tok/s est.
Runs on a Mac?
~9 tok/s est.

More Mistral models

Frequently asked questions

How much memory does Mistral Small 3.1 (Q8) need?

23.3 GB for the Q8_0 weights, plus ~4.0 GB of KV-cache at 16k context — about 27.3 GB total. Comfortable from 48 GB of VRAM or unified memory.

Does Mistral Small 3.1 (Q8) run on a Mac?

Yes — from the MacBook Pro M5 Pro 48GB (~9 tok/s est.). Unified memory means the RAM budget is the only limit.

What is the cheapest GPU for Mistral Small 3.1 (Q8)?

The AMD Ryzen AI Max+ 395 (Strix Halo) is the cheapest tracked card that runs Mistral Small 3.1 (Q8) comfortably — ~7 tok/s est. at ~$3,847 used (as of 2026-07-31).

Cite this page

ModelFit: Mistral Small 3.1 (Q8) — specs, memory math and hardware verdicts.
https://modelfit.io/models/mistral-small-3.1-24b/ (dataset updated 2026-09-03, CC BY 4.0).