GPT-OSS 120B
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~51 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for reasoning, coding, agents on RTX PRO 6000.
The RTX PRO 6000 Blackwell is NVIDIA's workstation flagship, built for professionals rather than gamers. Its 96GB of ECC GDDR7 is three times the RTX 5090's VRAM, at the same 1,792 GB/s bandwidth. That combination lets it load 120B-class dense models entirely in memory while still generating 7B-8B model tokens as fast as the 5090.
The best local LLM for the RTX PRO 6000 is GPT-OSS 120B at ~51 tok/s on its 96GB VRAM. It uses ~65.4GB of VRAM; the RTX PRO 6000 handles up to 120B parameter models at Q4. A 27B-class model (Qwen3.8 27B) at Q4 runs at ~45 tok/s.
Sizing rule: a Q4 model needs about 0.6 GB of VRAM per billion parameters, and ModelFit budgets 90% of this card's 96GB for weights, context, and KV-cache. The per-size table below uses that same budget. Other strong fits: Qwen3-Next 80B-A3B (80B, ~50.4GB) and Qwen3.5 122B-A10B Instruct (122B, ~72GB). GPT-OSS 120B runs at an estimated 51 tok/s on this card. What it will not run: dense models above ~120B at Q4 exceed this budget and need a lower quant, CPU offload, or a second GPU.
Speeds are ModelFit estimates from memory bandwidth and model size, not measured benchmarks.
Cite this page: ModelFit, RTX PRO 6000 Blackwell 96GB Local LLM (2026): 120B at ~51 tok/s, https://modelfit.io/gpu/rtx-6000-pro/, updated August 2026, CC BY 4.0.
Last updated: August 16, 2026 · Editor: ModelFit Team
*Launch MSRP was $8,565; workstation channel pricing has since risen with AI GPU demand
| Model Size | Est. Speed | Fit on 96GB |
|---|---|---|
| 7B | ~162 tok/s | Fits in VRAM |
| 14B | ~90 tok/s | Fits in VRAM |
| 20B MoE (3.6B active) | ~138 tok/s | Fits in VRAM |
| 32B | ~45 tok/s | Fits in VRAM |
| 35B MoE (3B active) | ~117 tok/s | Fits in VRAM |
| 70B | ~23 tok/s | Fits in VRAM |
| 120B MoE (5.1B active) | ~56 tok/s | Fits in VRAM |

ModelFit estimates, not measured benchmarks: anchored to an 8B-class Q4_K_M model at 16K context on the RTX PRO 6000's 1792 GB/s bandwidth, then scaled by model size. MoE rows scale by active parameters (decode reads only the active experts), so a 35B MoE runs far faster than a dense 32B. "CPU offload" sizes exceed the 96GB VRAM; dense models slow to a crawl there, MoE models degrade less because hot experts stay GPU-resident.
Context costs VRAM too. GPT-OSS 120B loads ~65.4 GB of weights; at 16k context the KV cache adds ~6.0 GB (still fits the ~86 GB usable VRAM), and at 64k it adds ~24.0 GB (exceeds the budget, use a smaller quant or a q8_0 KV cache).
KV-cache figures assume an fp16 cache, the llama.cpp/Ollama default. Standard GQA models use a size-class estimate (8 KV heads x 128 head dim class); hybrid linear-attention models (Qwen3.5/3.6, Qwen3-Next) use the exact per-token cost from their published config, since only their sparse full-attention layers cache KV. A q8_0 KV cache roughly halves either figure. Estimates, not measurements.
A Gen4 M.2 drive keeps your whole GGUF and quant collection on fast local storage, loading models straight off NVMe.
Check price on Amazon40Gbps external storage fast enough to run models from. Pair it with an M.2 drive for a portable model vault.
Check price on AmazonModelFit may earn a commission on purchases through these links, at no extra cost to you. Prices shown are approximate street references.
96GB of ECC GDDR7 at 1,792 GB/s gives the RTX PRO 6000 the same per-token speed as the RTX 5090 on models that fit both cards, since throughput is bandwidth-bound and the two cards share an identical bandwidth spec. The difference is capacity: a 70B model at Q4 (~42GB) leaves over 40GB free, and models up to roughly 120B parameters at Q4 fit inside the usable budget with room for a long context window. That headroom also means the card can hold two or three mid-size models resident at once, useful for running a chat model alongside a coding model without reloading. The 600W power draw and workstation pricing put this card well outside consumer territory; it is aimed at AI builders and studios that need the largest local models on a single GPU.
| Hardware | Memory | Speed | Bandwidth | Price |
|---|---|---|---|---|
| RTX 4090 | 24 GB | 104 tok/s | 1008 GB/s | $3,494 |
| Ryzen AI Max+ 395 | 110 GB | 30 tok/s | 256 GB/s | $3,847 |
| RTX 5090 | 32 GB | 145 tok/s | 1792 GB/s | $4,700 |
| RTX PRO 6000 | 96 GB | 145 tok/s | 1792 GB/s | $12,912 |
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~51 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for reasoning, coding, agents on RTX PRO 6000.
Qwen / 80B / Q4_K_M / ~50.4 GB
Best for: Chat, Coding, Long Context·Pop: 80/100
Perf: ~83 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for chat, coding, long context on RTX PRO 6000.
Qwen / 122B / Q4_K_M / ~72 GB
Best for: Frontier-level reasoning, Complex tasks·Pop: 75/100
Perf: ~41 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for frontier-level reasoning, complex tasks on RTX PRO 6000.
Llama / 109B / Q4_K_M / ~67 GB
Best for: Long context, Quality, Multimodal·Pop: 86/100
Perf: ~35 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for long context, quality, multimodal on RTX PRO 6000.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~73 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for reasoning, coding, agents on RTX PRO 6000.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agent scenarios·Pop: 90/100
Perf: ~73 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for reasoning, coding, agent scenarios on RTX PRO 6000.
Gemma / 26B / Q8_0 / ~28.1 GB
Best for: Chat, Coding, Multimodal·Pop: 86/100
Perf: ~73 tok/s · first token ~0.4s
Fits in 96 GB VRAM with room to spare. Best for chat, coding, multimodal on RTX PRO 6000.
Qwen / 27B / Q8_0 / ~30 GB
Best for: Coding, Quality, Long context·Pop: 92/100
Perf: ~32 tok/s · first token ~0.5s
Fits in 96 GB VRAM with room to spare. Best for coding, quality, long context on RTX PRO 6000.
Qwen / 80B / Q8_0 / ~84.8 GB
Best for: Chat, Coding, Long Context·Pop: 80/100
Perf: ~51 tok/s · first token ~1.0s
Fits in 96 GB VRAM with room to spare. Best for chat, coding, long context on RTX PRO 6000.
Llama / 70B / Q6_K / ~57.9 GB
Best for: Quality, Coding·Pop: 82/100
Perf: ~17 tok/s · first token ~1.2s
Fits in 96 GB VRAM with room to spare. Best for quality, coding on RTX PRO 6000.
The RTX PRO 6000 tops out around up to 120b parameter models. For anything bigger, an hourly rented GPU runs the same open weights with the same Ollama workflow, billed by the hour, no hardware purchase needed.
ModelFit may earn a commission on sign-ups made through these links, at no extra cost to you.
Alibaba Cloud: Widest size range (0.5B to 235B)
LlamaMeta: Most popular open-weight model family
DeepSeekDeepSeek AI: Best-in-class reasoning with R1 models
MistralMistral AI: Excellent performance-per-parameter ratio
GemmaGoogle DeepMind: Excellent quality at small sizes (1B-9B)
PhiMicrosoft: Best quality-per-gigabyte at small sizes
The RTX PRO 6000 Blackwell has 96GB of ECC GDDR7 VRAM with 1,792 GB/s bandwidth (NVIDIA, 2026). About 86GB is usable for model loading after driver overhead, enough for dense models up to roughly 120B parameters at Q4 quantization.
Up to roughly 120B parameter models at Q4 quantization fit inside its 96GB VRAM. Smaller 7B-32B models run with a large context window and headroom to keep a second model loaded at the same time.
It is a professional workstation card, not a consumer buy. Trading well above its $8,565 launch MSRP, it costs several times an RTX 5090, but it is one of the few single GPUs that holds 120B-class models entirely in VRAM without splitting across multiple cards.
Both share the same 1,792 GB/s bandwidth, so 7B-8B token speed is nearly identical (~145 tok/s). The RTX PRO 6000 has three times the VRAM (96GB vs 32GB), letting it hold much larger models, but it costs several times more and is not built for gaming.
No. It uses the same NVIDIA CUDA driver stack as GeForce cards, so Ollama detects it automatically. Keep drivers current, since workstation cards often get certified driver updates on a slightly different cadence than GeForce.
ModelFit estimates a 32B model on the RTX PRO 6000 runs at roughly 45 tok/s at Q4_K_M. The current 27B-class pick in the catalog is Qwen3.8 27B (ollama run qwen3.8:27b).
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Use our interactive wizard to compare models across Apple Silicon and NVIDIA GPUs.