GPU Value Index for Local LLMs

Estimated tokens/sec per dollar, every tracked card ranked. Open methodology, CC BY 4.0 data, dated prices.

Last updated: August 16, 2026 · Editor: ModelFit Team

Key facts
  • Best value overall: AMD Radeon RX 7900 XT at ~151 tok/s per $1,000 (used market).
  • Best value new NVIDIA: NVIDIA GeForce RTX 5070 Ti at ~84 tok/s per $1,000.
  • Fastest regardless of price: NVIDIA GeForce RTX 5090 at ~162 tok/s est. (8B Q4).

Cite this page: ModelFit, GPU Value Index for Local LLMs, https://modelfit.io/gpu/value-index/, updated 2026-08-16, CC BY 4.0. Speed figures are engine estimates, not measurements.

The ranking

GPUVRAMtok/s est.per $1kPrice
1RX 7900 XT20 GB83151~$550 used
2RX 7900 XTX24 GB100143~$700 used
3RTX 309024 GB97108~$900 used
4RTX 5070 Ti16 GB9784$1,152 (as of 2026-07-31)
5RTX 507012 GB6684$788 (as of 2026-07-31)
6RTX 5060 Ti16 GB5783$685 (as of 2026-07-31)
7RTX 306012 GB4778~$599 used (as of 2026-07-31)
8RTX 508016 GB10567$1,570 (as of 2026-07-31)
9RTX 4060 Ti16 GB3863~$599 used (as of 2026-07-31)
10RTX 407012 GB5861$945 (as of 2026-07-31)
11RTX 4070 SUPER12 GB6358~$1,089 used (as of 2026-07-31)
12RTX 40608 GB3457$599 (as of 2026-07-31)
13RTX 4070 Ti SUPER16 GB8155~$1,465 used (as of 2026-07-31)
14RTX 4080 SUPER16 GB8855~$1,600 used (as of 2026-07-31)
15RTX 509032 GB16234~$4,700 used (as of 2026-07-31)
16RTX 409024 GB11733~$3,494 used (as of 2026-07-31)
17RTX PRO 600096 GB16213~$12,912 used (as of 2026-07-31)
18Ryzen AI Max+ 395full system110 GB349~$3,847 used (as of 2026-07-31)

Method: tokens/sec are ModelFit engine estimates for an 8B Q4_K_M model at 16k context (bandwidth-bound, not measured). Value = est. tok/s / dated median listing price x 1,000. AMD scores assume a working ROCm setup. The unified-memory row is a complete system, not a card. Prices: see each card page for the live listing link.

FAQ

What is the best GPU value for local LLMs in 2026?

On estimated speed per dollar, the AMD Radeon RX 7900 XT leads this index at ~151 tokens/sec per $1,000 (83 tok/s est. on an 8B Q4 model at ~$550 used). That is a used-market price; among new NVIDIA cards, the NVIDIA GeForce RTX 5070 Ti leads at ~84 tok/s per $1,000.

Is the RTX 3090 still worth buying for local LLMs?

On value, yes: 24 GB of VRAM at ~$900 used gives it one of the best tokens-per-dollar scores among NVIDIA cards (~108 tok/s per $1,000). Newer 16 GB cards are faster per watt but hold far less of a model in VRAM.

Are AMD cards good for running local LLMs?

On paper they top the value table, but the estimate assumes a working ROCm setup (Linux, or Windows via recent drivers). If you want zero-friction Ollama, NVIDIA remains the safer buy; if you enjoy tuning, the RX 7900 series offers the most VRAM per dollar.

How is this index computed?

Tokens/sec figures are ModelFit engine estimates for an 8B Q4_K_M model at 16k context (bandwidth-bound model, not measurements). Value = estimated tokens/sec divided by the dated median listing price, scaled per $1,000. Prices are medians of legitimate listings, refreshed periodically; used-market cards are labeled. Data: CC BY 4.0.