Qwen3-Next 80B-A3B

Qwen3-Next 80B-A3B loads in 50.4 GB at Q4_K_M and is comfortable from 96 GB of memory, Macs included from the MacBook Pro M5 Max 128GB. Below: the full memory math, the cheapest card that runs it well, and the fastest machine we track.

PARAMETERS
80B (3B active)
FORMAT
Q4_K_M
MIN MEMORY
72 GB
BEST FOR
Chat, Coding, Long Context

Memory math

Weights (Q4_K_M)
50.4 GB
+ KV cache (16k)
~0.4 GB
Total at 16k
~50.8 GB
Comfortable from
96 GB

KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).

Hardware snapshot

Cheapest GPU
~17 tok/s est. — ~$3,847 used (as of 2026-07-31)
Fastest
~83 tok/s est.
Runs on a Mac?
~44 tok/s est.

Go deeper

Every Qwen3-Next 80B-A3B quant by real GGUF file sizeAlso tracked: Qwen3-Next 80B-A3B (Q8) (Q8_0, 84.8 GB, min 128 GB)

More Qwen models

Frequently asked questions

How much memory does Qwen3-Next 80B-A3B need?

50.4 GB for the Q4_K_M weights, plus ~0.4 GB of KV-cache at 16k context — about 50.8 GB total. Comfortable from 96 GB of VRAM or unified memory.

Does Qwen3-Next 80B-A3B run on a Mac?

Yes — from the MacBook Pro M5 Max 128GB (~44 tok/s est.). Unified memory means the RAM budget is the only limit.

What is the cheapest GPU for Qwen3-Next 80B-A3B?

The AMD Ryzen AI Max+ 395 (Strix Halo) is the cheapest tracked card that runs Qwen3-Next 80B-A3B comfortably — ~17 tok/s est. at ~$3,847 used (as of 2026-07-31).

Cite this page

ModelFit: Qwen3-Next 80B-A3B — specs, memory math and hardware verdicts.
https://modelfit.io/models/qwen3-next-80b-a3b/ (dataset updated 2026-09-03, CC BY 4.0).