Best hardware for GPT-OSS 120B
117B params, 5.1B active at MXFP4 — every tracked GPU and current Mac, graded. Engine estimates, dataset updated 2026-09-03.
GPT-OSS 120B on every tracked GPU
| GPU | Memory | Verdict | Est. speed | Value | |
|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell | 96 GB VRAM | Runs well | ~51 tok/s | 4 / $1k | |
| AMD Ryzen AI Max+ 395 (Strix Halo) | 110 GB unified memory | Runs well | ~11 tok/s | 3 / $1k | |
| NVIDIA GeForce RTX 5090 | 32 GB VRAM | No | ~18 tok/s | — | |
| NVIDIA GeForce RTX 4090 | 24 GB VRAM | No | ~13 tok/s | — | |
| NVIDIA GeForce RTX 5080 | 16 GB VRAM | No | ~12 tok/s | — | |
| AMD Radeon RX 7900 XTX | 24 GB VRAM | No | ~11 tok/s | — | |
| NVIDIA GeForce RTX 5070 Ti | 16 GB VRAM | No | ~11 tok/s | — | |
| NVIDIA GeForce RTX 3090 | 24 GB VRAM | No | ~11 tok/s | — | |
| NVIDIA GeForce RTX 4080 SUPER | 16 GB VRAM | No | ~10 tok/s | — | |
| AMD Radeon RX 7900 XT | 20 GB VRAM | No | ~9 tok/s | — | |
| NVIDIA GeForce RTX 4070 Ti SUPER | 16 GB VRAM | No | ~9 tok/s | — | |
| NVIDIA GeForce RTX 5070 | 12 GB VRAM | No | ~7 tok/s | — | |
| NVIDIA GeForce RTX 4070 SUPER | 12 GB VRAM | No | ~7 tok/s | — | |
| NVIDIA GeForce RTX 4070 | 12 GB VRAM | No | ~6 tok/s | — | |
| NVIDIA GeForce RTX 5060 Ti | 16 GB VRAM | No | ~6 tok/s | — | |
| NVIDIA GeForce RTX 3060 | 12 GB VRAM | No | ~5 tok/s | — | |
| NVIDIA GeForce RTX 4060 Ti | 16 GB VRAM | No | ~4 tok/s | — | |
| NVIDIA GeForce RTX 4060 | 8 GB VRAM | No | ~4 tok/s | — |
Value = est. tok/s per $1,000 at the dated median listing price. Cheapest: AMD Ryzen AI Max+ 395 (Strix Halo) (~$3,847 used (as of 2026-07-31)).
GPT-OSS 120B on current Macs (M5/M6)
| Mac | Memory | Verdict | Est. speed | |
|---|---|---|---|---|
| Mac Mini M6 16GB | 16 GB unified memory | No | ~3 tok/s | |
| MacBook Air M5 24GB | 24 GB unified memory | No | ~3 tok/s | |
| MacBook Pro M5 Pro 48GB | 48 GB unified memory | No | ~8 tok/s | |
| MacBook Pro M5 Max 128GB | 128 GB unified memory | Runs well | ~29 tok/s | |
| Mac Studio M5 Ultra 256GB | 256 GB unified memory | Runs well | ~43 tok/s |
Unified memory: the RAM budget (~70-85%) is the limit, not a VRAM wall. Current-gen configs only.
Frequently asked questions
What is the cheapest GPU that runs GPT-OSS 120B well?
The AMD Ryzen AI Max+ 395 (Strix Halo) is the cheapest tracked card where GPT-OSS 120B (MXFP4) runs comfortably — ~11 tok/s estimated, at ~$3,847 used (as of 2026-07-31).
What is the best value GPU for GPT-OSS 120B?
On estimated tokens per dollar, the NVIDIA RTX PRO 6000 Blackwell leads for GPT-OSS 120B at ~4 tok/s per $1,000 (~$12,912 used (as of 2026-07-31)). Prices are dated listing medians.
Does GPT-OSS 120B run on a Mac?
Yes. GPT-OSS 120B runs from the MacBook Pro M5 Max 128GB (~29 tok/s est.) up; unified memory means the RAM budget, not a VRAM wall, is the limit. Apple Silicon is often the cheapest path to large models.
What is the fastest way to run GPT-OSS 120B?
The fastest tracked machine for GPT-OSS 120B is the NVIDIA RTX PRO 6000 Blackwell at ~51 tok/s estimated. Speed is memory-bandwidth-bound, so high-bandwidth cards win.
Cite this page
ModelFit: Best hardware for GPT-OSS 120B (117B, MXFP4). https://modelfit.io/best-hardware-for/gpt-oss-120b/ (dataset updated 2026-09-03, CC BY 4.0).