Can you run Qwen3.8 27B on RTX 5090?
Qwen3.8 27B Q4_K_M on the NVIDIA GeForce RTX 5090 — verdict, VRAM math and estimated speed.
Yes, the NVIDIA GeForce RTX 5090 runs Qwen3.8 27B. 16.5 GB weights at Q4_K_M vs 28.8 GB usable VRAM; ~52 tok/s est..
VRAM math and speed are ModelFit engine estimates, not measurements. Commands are registry-verified Ollama tags.
Cite this page: ModelFit, Qwen3.8 27B on RTX 5090, https://modelfit.io/can-i-run/qwen3.8-27b-q4-on-rtx-5090/, updated August 2026, CC BY 4.0.
Last updated: August 16, 2026 · Editor: ModelFit Team
Memory math: weights + context vs budget
Weights take 16.5 GB. Context costs extra KV-cache on top — this is where long-context sessions break on cards that technically fit the weights.
KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.
Qwen3.8 27B on RTX 5090: FAQ
Can the NVIDIA GeForce RTX 5090 run Qwen3.8 27B?
Yes. Qwen3.8 27B (Q4_K_M) loads in about 16.5 GB and the RTX 5090 offers 28.8 GB of usable VRAM, leaving headroom for context. Expect at roughly 52 tokens/sec (est.).
How much VRAM does Qwen3.8 27B need?
About 16.5 GB for the weights at Q4_K_M, plus KV-cache for context: roughly 1.0 GB extra at 16k tokens. The RTX 5090 budget is 28.8 GB (32 GB x 90%).
What is the best quantization of Qwen3.8 27B for the RTX 5090?
Stick with the Q4_K_M build at 16.5 GB — every heavier quant exceeds the 28.8 GB usable VRAM.
What GPU do I need to run Qwen3.8 27B comfortably?
The RTX 5090 already runs Qwen3.8 27B comfortably. Larger cards only buy you longer context or a heavier quant.