Can you run Qwen3.5 35B-A3B Instruct on RTX 5080?

Qwen3.5 35B-A3B Instruct Q4_K_M on the NVIDIA GeForce RTX 5080: verdict, VRAM math and estimated speed.

Yes, but slow
Quick answer

Yes, but slowly: Qwen3.5 35B-A3B Instruct spills past the RTX 5080. 20 GB weights at Q4_K_M vs 14.4 GB usable VRAM; ~57 tok/s est. (ModelFit, 2026).

$ollama run qwen3.5:35b-a3b
VERDICT
Partial offload
EST. SPEED
~57 tok/s
WEIGHTS
20 GB Q4_K_M

VRAM math and speed are ModelFit engine estimates, not measurements. Commands are registry-verified Ollama tags.

Cite this page: ModelFit, Qwen3.5 35B-A3B Instruct on RTX 5080, https://modelfit.io/can-i-run/qwen3.5-35b-a3b-q4-on-rtx-5080/, updated September 2026, CC BY 4.0.

Last updated: September 18, 2026 · Editor: ModelFit Team

VRAM
16 GB (14.4 usable)
Model weights
20 GB Q4_K_M
Est. speed
~57 tok/s
First token
~1s
Fit grade
C · 0

What limits this combo

The four constraints the fit engine checks for Qwen3.5 35B-A3B Instruct on the RTX 5080, in order of what usually breaks first.

Weights fitBlocked20 GB of weights against 14.4 GB usable — the model does not load fully in VRAM; the shortfall spills to system RAM.
Context ceilingBlockedThe weights alone exceed the budget, so no usable context fits on top.
SpeedOK~57 tok/s (est.) on 960 GB/s of memory bandwidth — comfortable for interactive chat and agentic loops.
Use-case ceilingTightStrongest fit: Chat (C). Agentic coding caps at C — either the speed floor or the model's tuning is the limit; see the workload table for the per-case reason.

Quant explorer

Switch between the quality-gated builds tracked for this exact combo. Q4_K_M stays the default recommendation; heavier Q6/Q8 builds only appear when their exact Ollama tags are registry-verified.

ModelFit engine estimate
Verdict
Slow
Est. speed
~57 tok/s
First token
~1.0s · Instant
Safe context
n/a
C · 0 Tight fit

Why no 2-bit builds: Q1/Q2-class quants can look attractive in a memory table, but their quality loss is large enough that ModelFit excludes them from rankings instead of inflating the catalog with junk options. Usable VRAM: 14.4 GB.

Workload verdicts

Qwen3.5 35B-A3B Instruct on the RTX 5080, graded per use case from the model's registry-verified tuning and the engine's speed estimate for this exact combo.

ModelFit engine estimate
ChatC~57 tok/s est. vs ~20 needed for chat; partial offload slows everything.
CodingC~57 tok/s est. vs ~15 needed for coding; partial offload slows everything.
Agentic codingC~57 tok/s est. vs ~30 needed for agentic coding; partial offload slows everything.
ReasoningC~57 tok/s est. vs ~12 needed for reasoning; partial offload slows everything.
RAGC~57 tok/s est. vs ~20 needed for rag; 32k context does not fit; partial offload slows everything.

Grades combine tag-verified tuning (a model not built for the workload caps at C) with the tok/s floor each workload needs to feel usable. A combo that partially offloads caps at C; one that does not fit is D everywhere.

Memory math: weights + context vs budget

Weights take 20 GB. Context costs extra KV-cache on top — this is where long-context sessions break on cards that technically fit the weights.

Weights 20 GBKV cache · 16k context 0.3 GBRuntime reserve 1.6 GB

This configuration exceeds the comfortable VRAM budget by about 5.9 GB. Expect offload pressure or a shorter safe context.

ContextKV-cacheTotalFits
8k tokens0.2 GB20.2 GBOver
16k tokens0.3 GB20.3 GBOver
32k tokens0.6 GB20.6 GBOver
64k tokens1.3 GB21.3 GBOver
128k tokens2.5 GB22.5 GBOver

KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.

See how fast it feels

A deterministic typing simulation for Qwen3.5 35B-A3B Instruct: first token ~1.0s (instant prefill), then ~57 tokens/sec.

Start the simulation to preview the response pace.

ModelFit engine estimate, not a measured benchmark. Real speed varies with prompt length, thermals, runtime, and KV-cache settings.

Upgrade path

The cheapest tracked card that runs Qwen3.5 35B-A3B Instruct comfortably is the AMD Radeon RX 7900 XTX (24 GB VRAM).

See the RX 7900 XTX page

Qwen3.5 35B-A3B Instruct on RTX 5080: FAQ

Can the NVIDIA GeForce RTX 5080 run Qwen3.5 35B-A3B Instruct?

Barely. Qwen3.5 35B-A3B Instruct (Q4_K_M) needs about 20 GB but the RTX 5080 has 14.4 GB usable VRAM, so part of the model spills to system RAM and speed drops to at roughly 57 tokens/sec (est.).

How much VRAM does Qwen3.5 35B-A3B Instruct need?

About 20 GB for the weights at Q4_K_M, plus KV-cache for context: roughly 0.3 GB extra at 16k tokens. The RTX 5080 budget is 14.4 GB (16 GB x 90%).

What is the best quantization of Qwen3.5 35B-A3B Instruct for the RTX 5080?

Stick with the Q4_K_M build at 20 GB — every heavier quant exceeds the 14.4 GB usable VRAM.

What GPU do I need to run Qwen3.5 35B-A3B Instruct comfortably?

The cheapest tracked card that runs Qwen3.5 35B-A3B Instruct (Q4_K_M) fully in VRAM is the AMD Radeon RX 7900 XTX (24 GB).

How much context can Qwen3.5 35B-A3B Instruct use on the RTX 5080?

None worth having — the 20 GB of weights alone exceed the 14.4 GB usable VRAM, so the model only runs with system-RAM offload and short prompts.

Is Qwen3.5 35B-A3B Instruct on the RTX 5080 fast enough for coding agents?

Yes: the engine estimates ~57 tok/s against the ~30 tok/s an agentic loop needs to feel responsive, and Qwen3.5 35B-A3B Instruct is tuned for tool-calling.