Qwen3-Next 80B-A3B
Qwen / 80B / Q4_K_M / ~50.4 GB
Best for: Chat, Coding, Long Context·Perf: ~44.6 tok/s · first token ~1.5s
Best for chat, coding, long context. Strong fit for 96 GB RAM with balanced speed and quality.
Ranked open-weight models that run well on a 96GB machine, with estimated speed and the exact ollama command for each.
Try: RTX 4090, MacBook Pro M4, 16GB, iPhone 17 Pro
Estimates assume a representative Apple-Silicon machine with 96GB unified memory. Tok/s are ModelFit estimates, not measured benchmarks. Run the wizard for figures tuned to your exact chip.
Qwen / 80B / Q4_K_M / ~50.4 GB
Best for: Chat, Coding, Long Context·Perf: ~44.6 tok/s · first token ~1.5s
Best for chat, coding, long context. Strong fit for 96 GB RAM with balanced speed and quality.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agents·Perf: ~36.9 tok/s · first token ~1.5s
Best for reasoning, coding, agents. Strong fit for 96 GB RAM with balanced speed and quality.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agent scenarios·Perf: ~36.9 tok/s · first token ~1.5s
Best for reasoning, coding, agent scenarios. Strong fit for 96 GB RAM with balanced speed and quality.
Nemotron / 30B / Q4_K_M / ~23.7 GB
Best for: Agentic, Coding, Long context·Perf: ~72.8 tok/s · first token ~1.4s
Best for agentic, coding, long context. Strong fit for 96 GB RAM with balanced speed and quality.
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Perf: ~29.8 tok/s · first token ~1.6s
This model may feel memory-heavy on 96 GB RAM, but it is still listed for balanced speed and quality.
Qwen / 122B / Q4_K_M / ~72 GB
Best for: Frontier-level reasoning, Complex tasks·Perf: ~19 tok/s · first token ~1.8s
This model may feel memory-heavy on 96 GB RAM, but it is still listed for balanced speed and quality.
Llama / 109B / Q4_K_M / ~67 GB
Best for: Long context, Quality, Multimodal·Perf: ~16.1 tok/s · first token ~1.9s
This model may feel memory-heavy on 96 GB RAM, but it is still listed for balanced speed and quality.
Llama / 70B / Q4_K_M / ~42 GB
Best for: Quality, Coding·Perf: ~9.9 tok/s · first token ~2.3s
Best for quality, coding. Strong fit for 96 GB RAM with balanced speed and quality.
Gemma / 26B / Q8_0 / ~28.1 GB
Best for: Chat, Coding, Multimodal·Perf: ~37.1 tok/s · first token ~0.7s
Best for chat, coding, multimodal. Strong fit for 96 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Perf: ~67.4 tok/s · first token ~1.4s
Best for reasoning, coding, agents. Strong fit for 96 GB RAM with balanced speed and quality.
Qwen / 27B / Q8_0 / ~27.1 GB
Best for: Coding, Agent, Vision, Long context·Perf: ~14 tok/s · first token ~1.2s
Best for coding, agent, vision, long context. Strong fit for 96 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~20 GB
Best for: Reasoning, Coding, Agent scenarios·Perf: ~67.4 tok/s · first token ~1.4s
Best for reasoning, coding, agent scenarios. Strong fit for 96 GB RAM with balanced speed and quality.
The wizard tunes picks and speed estimates to your exact device, chip, and RAM.