Phi-4 Mini 3.8B

Phi — 3.8B params — Q4_K_M, local via Ollama. Dataset updated 2026-09-03.

PARAMETERS
3.8B
FORMAT
Q4_K_M
MIN MEMORY
6 GB
BEST FOR
Coding, Chat

Memory math

Weights (Q4_K_M)
3.2 GB
+ KV cache (16k)
~1.8 GB
Total at 16k
~5.0 GB
Comfortable from
8 GB

KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).

Hardware snapshot

Cheapest GPU
~139 tok/s est. — ~$550 used
Fastest
~273 tok/s est.
Runs on a Mac?
~59 tok/s est.

Go deeper

Every Phi-4 Mini 3.8B quant by real GGUF file sizeAlso tracked: Phi-4 Mini 3.8B (Q8) (Q8_0, 3.8 GB, min 8 GB)

More Phi models

Frequently asked questions

How much memory does Phi-4 Mini 3.8B need?

3.2 GB for the Q4_K_M weights, plus ~1.8 GB of KV-cache at 16k context — about 5.0 GB total. Comfortable from 8 GB of VRAM or unified memory.

Does Phi-4 Mini 3.8B run on a Mac?

Yes — from the Mac Mini M6 16GB (~59 tok/s est.). Unified memory means the RAM budget is the only limit.

What is the cheapest GPU for Phi-4 Mini 3.8B?

The AMD Radeon RX 7900 XT is the cheapest tracked card that runs Phi-4 Mini 3.8B comfortably — ~139 tok/s est. at ~$550 used.

Cite this page

ModelFit: Phi-4 Mini 3.8B — specs, memory math and hardware verdicts.
https://modelfit.io/models/phi4-mini-3.8b/ (dataset updated 2026-09-03, CC BY 4.0).