Qwen3.5 35B-A3B Instruct quants compared

28 GGUF builds by real file size, probed from bartowski/Qwen_Qwen3.5-35B-A3B-GGUF on Hugging Face (2026-09-02). 35B params, 3B active.

Download Qwen3.5 35B-A3B Instruct Q4_K_M (20.75 GB) — it fits 24 GB of memory with 16k context. With Ollama: ollama run qwen3.5:35b-a3b

Every Qwen3.5 35B-A3B Instruct quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1666.19 GB0.3 GB66.5 GB96 GBFull precision (lossless)
Q8_035.22 GB0.3 GB35.5 GB48 GBNear-lossless
Q6_K_L29.05 GB0.3 GB29.4 GB48 GBExcellent
Q6_K28.82 GB0.3 GB29.1 GB48 GBExcellent
Q5_K_L24.43 GB0.3 GB24.7 GB32 GBVery high
Q5_K_M24.13 GB0.3 GB24.4 GB32 GBVery high
Q5_K_S23.33 GB0.3 GB23.6 GB32 GBVery high
Q4_K_M *20.75 GB0.3 GB21.1 GB24 GBHigh — the default pick
Q4_121.3 GB0.3 GB21.6 GB32 GBHigh
Q4_K_L21.11 GB0.3 GB21.4 GB24 GBHigh
Q4_K_S20.01 GB0.3 GB20.3 GB24 GBHigh
Q4_019.41 GB0.3 GB19.7 GB24 GBHigh
IQ4_NL19.33 GB0.3 GB19.6 GB24 GBHigh
IQ4_XS18.35 GB0.3 GB18.7 GB24 GBHigh
Q3_K_XL16.97 GB0.3 GB17.3 GB24 GBAcceptable — visible loss
IQ3_M16.57 GB0.3 GB16.9 GB24 GBAcceptable — visible loss
Q3_K_L16.56 GB0.3 GB16.9 GB24 GBAcceptable — visible loss
IQ3_XS15.94 GB0.3 GB16.3 GB24 GBAcceptable — visible loss
Q3_K_M15.94 GB0.3 GB16.3 GB24 GBAcceptable — visible loss
Q3_K_S15.28 GB0.3 GB15.6 GB24 GBAcceptable — visible loss
IQ3_XXS14.68 GB0.3 GB15.0 GB24 GBAcceptable — visible loss
Q2_K_L13.04 GB0.3 GB13.4 GB16 GBExperimental — not ranked — never recommended
Q2_K12.58 GB0.3 GB12.9 GB16 GBExperimental — not ranked — never recommended
IQ2_M12.07 GB0.3 GB12.4 GB16 GBExperimental — not ranked — never recommended
IQ2_S11.09 GB0.3 GB11.4 GB16 GBExperimental — not ranked — never recommended
IQ2_XS10.89 GB0.3 GB11.2 GB16 GBExperimental — not ranked — never recommended
IQ2_XXS9.94 GB0.3 GB10.3 GB12 GBExperimental — not ranked — never recommended
IQ1_M8.77 GB0.3 GB9.1 GB12 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from bartowski/Qwen_Qwen3.5-35B-A3B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best Qwen3.5 35B-A3B Instruct quant by memory

MemoryRecommended quantTotal (16k ctx)
24 GBQ4_K_M21.1 GB
32 GBQ5_K_M24.4 GB
48 GBQ8_035.5 GB
96 GBBF1666.5 GB

Why we don't rank Qwen3.5 35B-A3B Instruct's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of Qwen3.5 35B-A3B Instruct?

Q4_K_M is the default pick: 20.75 GB of weights, high — the default pick quality, fitting comfortably in 24 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does Qwen3.5 35B-A3B Instruct need?

At Q4_K_M, Qwen3.5 35B-A3B Instruct needs 20.75 GB for the weights plus ~0.3 GB of KV-cache at 16k context — about 21.1 GB total, so a 24 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of Qwen3.5 35B-A3B Instruct?

No. Qwen3.5 35B-A3B Instruct at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: Qwen3.5 35B-A3B Instruct quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/qwen3.5-35b-a3b/ (data probed 2026-09-02, CC BY 4.0).