Can you run Qwen3.5 35B-A3B Instruct on RTX 5090?
Qwen3.5 35B-A3B Instruct Q4_K_M on the NVIDIA GeForce RTX 5090: verdict, VRAM math and estimated speed.
Yes, the NVIDIA GeForce RTX 5090 runs Qwen3.5 35B-A3B Instruct. 20 GB weights at Q4_K_M vs 28.8 GB usable VRAM; ~118 tok/s est. (ModelFit, 2026).
VRAM math and speed are ModelFit engine estimates, not measurements. Commands are registry-verified Ollama tags.
Cite this page: ModelFit, Qwen3.5 35B-A3B Instruct on RTX 5090, https://modelfit.io/can-i-run/qwen3.5-35b-a3b-q4-on-rtx-5090/, updated September 2026, CC BY 4.0.
Last updated: September 18, 2026 · Editor: ModelFit Team
What limits this combo
The four constraints the fit engine checks for Qwen3.5 35B-A3B Instruct on the RTX 5090, in order of what usually breaks first.
Quant explorer
Switch between the quality-gated builds tracked for this exact combo. Q4_K_M stays the default recommendation; heavier Q6/Q8 builds only appear when their exact Ollama tags are registry-verified.
Why no 2-bit builds: Q1/Q2-class quants can look attractive in a memory table, but their quality loss is large enough that ModelFit excludes them from rankings instead of inflating the catalog with junk options. Usable VRAM: 28.8 GB.
Workload verdicts
Qwen3.5 35B-A3B Instruct on the RTX 5090, graded per use case from the model's registry-verified tuning and the engine's speed estimate for this exact combo.
Grades combine tag-verified tuning (a model not built for the workload caps at C) with the tok/s floor each workload needs to feel usable. A combo that partially offloads caps at C; one that does not fit is D everywhere.
Memory math: weights + context vs budget
Weights take 20 GB. Context costs extra KV-cache on top — this is where long-context sessions break on cards that technically fit the weights.
The model leaves about 8.5 GB of the usable VRAM budget free after weights and 16k context.
KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.
See how fast it feels
A deterministic typing simulation for Qwen3.5 35B-A3B Instruct: first token ~0.9s (instant prefill), then ~118 tokens/sec.
ModelFit engine estimate, not a measured benchmark. Real speed varies with prompt length, thermals, runtime, and KV-cache settings.
Qwen3.5 35B-A3B Instruct on RTX 5090: FAQ
Can the NVIDIA GeForce RTX 5090 run Qwen3.5 35B-A3B Instruct?
Yes. Qwen3.5 35B-A3B Instruct (Q4_K_M) loads in about 20 GB and the RTX 5090 offers 28.8 GB of usable VRAM, leaving headroom for context. Expect at roughly 118 tokens/sec (est.).
How much VRAM does Qwen3.5 35B-A3B Instruct need?
About 20 GB for the weights at Q4_K_M, plus KV-cache for context: roughly 0.3 GB extra at 16k tokens. The RTX 5090 budget is 28.8 GB (32 GB x 90%).
What is the best quantization of Qwen3.5 35B-A3B Instruct for the RTX 5090?
Stick with the Q4_K_M build at 20 GB — every heavier quant exceeds the 28.8 GB usable VRAM.
What GPU do I need to run Qwen3.5 35B-A3B Instruct comfortably?
The RTX 5090 already runs Qwen3.5 35B-A3B Instruct comfortably. Larger cards only buy you longer context or a heavier quant.
How much context can Qwen3.5 35B-A3B Instruct use on the RTX 5090?
Up to 128k tokens stay fully in VRAM (22.5 GB total with the KV cache). Quantizing the KV cache to q8_0 roughly halves the KV column and buys back about one context tier.
Is Qwen3.5 35B-A3B Instruct on the RTX 5090 fast enough for coding agents?
Yes: the engine estimates ~118 tok/s against the ~30 tok/s an agentic loop needs to feel responsive, and Qwen3.5 35B-A3B Instruct is tuned for tool-calling.