Qwen3.6 35B-A3B
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~118 tok/s · first token ~0.9s
Fits in 32 GB VRAM with room to spare. Best for reasoning, coding, agents on RTX 5090.
The RTX 5090 redefines what is possible with consumer GPUs. With 32GB GDDR7 and an unprecedented ~145 tokens per second on 8B models, it can load 70B parameter models with Q4 quantization. The ultimate card for running the largest local AI models.
The best local LLM for the RTX 5090 is Qwen3.6 35B-A3B at ~118 tok/s on its 32GB VRAM. It uses ~22GB of VRAM; the RTX 5090 handles up to 32B parameter models at Q4. A 27B-class model (Qwen3.5 27B Instruct) runs at ~45 tok/s.
Speeds are ModelFit estimates from memory bandwidth and model size, not measured benchmarks.
*Launch MSRP was $1,999; current street pricing sits well above that amid the 2026 memory shortage
| Model Size | Est. Speed | Fit on 32GB |
|---|---|---|
| 7B | ~162 tok/s | Fits in VRAM |
| 14B | ~90 tok/s | Fits in VRAM |
| 20B MoE (3.6B active) | ~138 tok/s | Fits in VRAM |
| 32B | ~45 tok/s | Fits in VRAM |
| 35B MoE (3B active) | ~117 tok/s | Fits in VRAM |
| 70B | ~23 tok/s | Fits in VRAM |
| 120B MoE (5.1B active) | ~19 tok/s | CPU offload (slow) |
ModelFit estimates, not measured benchmarks: anchored to an 8B-class Q4_K_M model at 16K context on the RTX 5090's 1792 GB/s bandwidth, then scaled by model size. MoE rows scale by active parameters (decode reads only the active experts), so a 35B MoE runs far faster than a dense 32B. "CPU offload" sizes exceed the 32GB VRAM; dense models slow to a crawl there, MoE models degrade less because hot experts stay GPU-resident.
Context costs VRAM too. Qwen3.6 35B-A3B loads ~22 GB of weights; at 16k context the KV cache adds ~0.3 GB (still fits the ~29 GB usable VRAM), and at 64k it adds ~1.3 GB (still fits).
KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.
A Gen4 M.2 drive keeps your whole GGUF and quant collection on fast local storage, loading models straight off NVMe.
Check price on Amazon40Gbps external storage fast enough to run models from. Pair it with an M.2 drive for a portable model vault.
Check price on AmazonModelFit may earn a commission on purchases through these links, at no extra cost to you. Prices shown are approximate street references.
32GB GDDR7 at 1,792 GB/s is unprecedented in a consumer card. This is enough to load 32B models at Q6 or even Q8 quantization with room to spare, delivering higher quality than Q4 on smaller cards. The 5090 also handles multi-model setups: run a 14B chat model alongside a 7B coding model simultaneously. For the largest 70B models at Q3_K_M (~30GB), it fits with about 1GB headroom. For AI enthusiasts who want to run the biggest models locally, this is the endgame card.
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~118 tok/s · first token ~0.9s
Fits in 32 GB VRAM with room to spare. Best for reasoning, coding, agents on RTX 5090.
Qwen / 35B / Q4_K_M / ~20 GB
Best for: Reasoning, Coding, Agent scenarios·Pop: 90/100
Perf: ~118 tok/s · first token ~0.9s
Fits in 32 GB VRAM with room to spare. Best for reasoning, coding, agent scenarios on RTX 5090.
Gemma / 26B / Q4_K_M / ~16 GB
Best for: Chat, Coding, Multimodal·Pop: 86/100
Perf: ~118 tok/s · first token ~0.3s
Fits in 32 GB VRAM with room to spare. Best for chat, coding, multimodal on RTX 5090.
Qwen / 27B / Q4_K_M / ~16 GB
Best for: Chat, Coding, Complex reasoning·Pop: 82/100
Perf: ~52 tok/s · first token ~0.4s
Fits in 32 GB VRAM with room to spare. Best for chat, coding, complex reasoning on RTX 5090.
Qwen / 27B / Q4_K_M / ~18 GB
Best for: Coding, Quality, Long context·Pop: 92/100
Perf: ~52 tok/s · first token ~0.4s
Fits in 32 GB VRAM with room to spare. Best for coding, quality, long context on RTX 5090.
Qwen / 30B / Q4_K_M / ~22 GB
Best for: Quality, Coding·Pop: 78/100
Perf: ~125 tok/s · first token ~0.9s
Fits in 32 GB VRAM with room to spare. Best for quality, coding on RTX 5090.
Gemma / 31B / Q4_K_M / ~20 GB
Best for: Quality, Coding, Multimodal·Pop: 84/100
Perf: ~46 tok/s · first token ~1.0s
Fits in 32 GB VRAM with room to spare. Best for quality, coding, multimodal on RTX 5090.
Qwen / 14B / Q8_0 / ~15.9 GB
Best for: Coding, Quality·Pop: 84/100
Perf: ~56 tok/s · first token ~0.4s
Fits in 32 GB VRAM with room to spare. Best for coding, quality on RTX 5090.
Nemotron / 30B / Q6_K / ~24 GB
Best for: Reasoning, Math, Agentic tasks·Pop: 60/100
Perf: ~93 tok/s · first token ~1.0s
Fits in 32 GB VRAM with room to spare. Best for reasoning, math, agentic tasks on RTX 5090.
Gemma / 26B / Q8_0 / ~28.1 GB
Best for: Chat, Coding, Multimodal·Pop: 86/100
Perf: ~73 tok/s · first token ~0.4s
Fits in 32 GB VRAM with room to spare. Best for chat, coding, multimodal on RTX 5090.
The RTX 5090 tops out around up to 70b parameter models. For anything bigger, an hourly rented GPU runs the same open weights with the same Ollama workflow, billed by the hour, no hardware purchase needed.
ModelFit may earn a commission on sign-ups made through these links, at no extra cost to you.
Alibaba Cloud: Widest size range (0.5B to 235B)
LlamaMeta: Most popular open-weight model family
DeepSeekDeepSeek AI: Best-in-class reasoning with R1 models
MistralMistral AI: Excellent performance-per-parameter ratio
GemmaGoogle DeepMind: Excellent quality at small sizes (1B-9B)
PhiMicrosoft: Best quality-per-gigabyte at small sizes
The RTX 5090 has 32GB GDDR7 VRAM with 1,792 GB/s bandwidth, the most of any consumer GPU. About 31GB is usable for models. This fits 70B models at Q3 quantization and 32B models at Q8 for maximum quality.
Up to 70B parameter models at Q3 quantization, or 32B models at Q8 for best quality. The 5090 is the only consumer GPU that can run Llama 3.1 70B locally. For 7B-8B models it delivers an incredible ~145 tok/s (14B lands around an estimated 90).
The RTX 5090 is the absolute best consumer GPU for local AI. At ~145 tok/s on 8B models, responses feel instant. Its 32GB VRAM unlocks 70B models that no other consumer card can handle. The $2,499 price is justified if you need large models.
The RTX 5090 is 39% faster on 8B models (an estimated 145 vs 104 tok/s) with 33% more VRAM (32 vs 24GB). It runs 70B models that the 4090 cannot. At similar pricing (~$2,500), the 5090 is the clear choice for new purchases in 2026.
For many use cases, yes. A 32B model at Q6 on the 5090 rivals GPT-4 quality for coding and reasoning tasks, running at an estimated 30-35 tok/s with low latency and complete privacy. The main gap is context length: cloud models support 128K+ tokens while local 70B models are limited to 8K-16K.
ModelFit estimates a 32B model on the RTX 5090 runs at roughly 45 tok/s at Q4_K_M. The current 27B-class pick in the catalog is Qwen3.5 27B Instruct (ollama run qwen3.5:27b).
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Use our interactive wizard to compare models across Apple Silicon and NVIDIA GPUs.