Can you run Qwen3.8 27B on RTX 5090?

Qwen3.8 27B Q4_K_M on the NVIDIA GeForce RTX 5090 — verdict, VRAM math and estimated speed.

Yes, it runs
Quick answer

Yes, the NVIDIA GeForce RTX 5090 runs Qwen3.8 27B. 16.5 GB weights at Q4_K_M vs 28.8 GB usable VRAM; ~52 tok/s est..

$ollama run qwen3.8:27b
VERDICT
Comfortable
EST. SPEED
~52 tok/s
WEIGHTS
16.5 GB Q4_K_M

VRAM math and speed are ModelFit engine estimates, not measurements. Commands are registry-verified Ollama tags.

Cite this page: ModelFit, Qwen3.8 27B on RTX 5090, https://modelfit.io/can-i-run/qwen3.8-27b-q4-on-rtx-5090/, updated September 2026, CC BY 4.0.

Last updated: September 3, 2026 · Editor: ModelFit Team

VRAM
32 GB (28.8 usable)
Model weights
16.5 GB Q4_K_M
Est. speed
~52 tok/s
First token
~0.4s
Fit grade
A · 43

Quant explorer

Switch between the quality-gated builds tracked for this exact combo. Q4_K_M stays the default recommendation; heavier Q6/Q8 builds only appear when their exact Ollama tags are registry-verified.

ModelFit engine estimate
Verdict
Runs
Est. speed
~52 tok/s
First token
~0.4s · Instant
Safe context
128k
A · 43 Excellent headroom

Why no 2-bit builds: Q1/Q2-class quants can look attractive in a memory table, but their quality loss is large enough that ModelFit excludes them from rankings instead of inflating the catalog with junk options. Usable VRAM: 28.8 GB.

Workload verdicts

Qwen3.8 27B on the RTX 5090, graded per use case from the model's registry-verified tuning and the engine's speed estimate for this exact combo.

ModelFit engine estimate
ChatA~52 tok/s est. vs ~20 needed for chat.
CodingA~52 tok/s est. vs ~15 needed for coding.
Agentic codingA~52 tok/s est. vs ~30 needed for agentic coding.
ReasoningC~52 tok/s est. vs ~12 needed for reasoning; not a reasoning-tuned model.
RAGA~52 tok/s est. vs ~20 needed for rag.

Grades combine tag-verified tuning (a model not built for the workload caps at C) with the tok/s floor each workload needs to feel usable. A combo that partially offloads caps at C; one that does not fit is D everywhere.

Memory math: weights + context vs budget

Weights take 16.5 GB. Context costs extra KV-cache on top — this is where long-context sessions break on cards that technically fit the weights.

Weights 17 GBKV cache · 16k context 1.0 GBHeadroom 11 GBRuntime reserve 3.2 GB

The model leaves about 11 GB of the usable VRAM budget free after weights and 16k context.

ContextKV-cacheTotalFits
8k tokens0.5 GB17.0 GBFits
16k tokens1.0 GB17.5 GBFits
32k tokens2.0 GB18.5 GBFits
64k tokens4.0 GB20.5 GBFits
128k tokens8.0 GB24.5 GBFits

KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.

See how fast it feels

A deterministic typing simulation for Qwen3.8 27B: first token ~0.4s (instant prefill), then ~52 tokens/sec.

Start the simulation to preview the response pace.

ModelFit engine estimate, not a measured benchmark. Real speed varies with prompt length, thermals, runtime, and KV-cache settings.

Qwen3.8 27B on RTX 5090: FAQ

Can the NVIDIA GeForce RTX 5090 run Qwen3.8 27B?

Yes. Qwen3.8 27B (Q4_K_M) loads in about 16.5 GB and the RTX 5090 offers 28.8 GB of usable VRAM, leaving headroom for context. Expect at roughly 52 tokens/sec (est.).

How much VRAM does Qwen3.8 27B need?

About 16.5 GB for the weights at Q4_K_M, plus KV-cache for context: roughly 1.0 GB extra at 16k tokens. The RTX 5090 budget is 28.8 GB (32 GB x 90%).

What is the best quantization of Qwen3.8 27B for the RTX 5090?

The Q8_0 build is the highest quality that fits (27.1 GB vs 28.8 GB usable). The Q4_K_M build at 16.5 GB leaves more room for long context.

What GPU do I need to run Qwen3.8 27B comfortably?

The RTX 5090 already runs Qwen3.8 27B comfortably. Larger cards only buy you longer context or a heavier quant.