Qwen3.8-Flash-Next

Qwen3.8-Flash-Next loads in 123 GB at IQ1_S and is comfortable from 192 GB of memory, Macs included from the Mac Studio M5 Ultra 256GB. Below: the full memory math, the cheapest card that runs it well, and the fastest machine we track.

PARAMETERS
125B (6B active)
FORMAT
IQ1_S
MIN MEMORY
192 GB
BEST FOR
Agentic coding, Reasoning, Multimodal

Memory math

Weights (IQ1_S)
123 GB
+ KV cache (16k)
~6.0 GB
Total at 16k
~129.0 GB
Comfortable from
192 GB

KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).

Hardware snapshot

Cheapest GPU
no consumer fit
Fastest
~32 tok/s est.
Runs on a Mac?
~32 tok/s est.

More Qwen models

Frequently asked questions

How much memory does Qwen3.8-Flash-Next need?

123 GB for the IQ1_S weights, plus ~6.0 GB of KV-cache at 16k context — about 129.0 GB total. Comfortable from 192 GB of VRAM or unified memory.

Does Qwen3.8-Flash-Next run on a Mac?

Yes — from the Mac Studio M5 Ultra 256GB (~32 tok/s est.). Unified memory means the RAM budget is the only limit.

What is the cheapest GPU for Qwen3.8-Flash-Next?

No tracked consumer GPU runs Qwen3.8-Flash-Next comfortably.

Cite this page

ModelFit: Qwen3.8-Flash-Next — specs, memory math and hardware verdicts.
https://modelfit.io/models/qwen3.8-flash-next/ (dataset updated 2026-09-03, CC BY 4.0).