GPT-OSS 120B
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Perf: ~42.8 tok/s · first token ~1.5s
Best for reasoning, coding, agents. Strong fit for 128 GB RAM with balanced speed and quality.
Ranked open-weight models that run well on a 128GB machine, with estimated speed and the exact ollama command for each.
Try: RTX 4090, MacBook Pro M4, 16GB, iPhone 17 Pro
Estimates assume a representative Apple-Silicon machine with 128GB unified memory. Tok/s are ModelFit estimates, not measured benchmarks. Run the wizard for figures tuned to your exact chip.
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Perf: ~42.8 tok/s · first token ~1.5s
Best for reasoning, coding, agents. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 122B / Q4_K_M / ~72 GB
Best for: Frontier-level reasoning, Complex tasks·Perf: ~28.4 tok/s · first token ~1.6s
Best for frontier-level reasoning, complex tasks. Strong fit for 128 GB RAM with balanced speed and quality.
Llama / 109B / Q4_K_M / ~67 GB
Best for: Long context, Quality, Multimodal·Perf: ~23.1 tok/s · first token ~1.7s
Best for long context, quality, multimodal. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 80B / Q8_0 / ~84.8 GB
Best for: Chat, Coding, Long Context·Perf: ~35 tok/s · first token ~1.5s
This model may feel memory-heavy on 128 GB RAM, but it is still listed for balanced speed and quality.
Qwen / 80B / Q4_K_M / ~50.4 GB
Best for: Chat, Coding, Long Context·Perf: ~64 tok/s · first token ~1.4s
Best for chat, coding, long context. Strong fit for 128 GB RAM with balanced speed and quality.
Laguna / 118B / Q4_K_M / ~96 GB
Best for: Agentic coding, Long-horizon tasks·Perf: ~32.3 tok/s · first token ~1.6s
This model may feel memory-heavy on 128 GB RAM, but it is still listed for balanced speed and quality.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agents·Perf: ~53 tok/s · first token ~1.4s
Best for reasoning, coding, agents. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 35B / Q8_0 / ~38.7 GB
Best for: Reasoning, Coding, Agent scenarios·Perf: ~53 tok/s · first token ~1.4s
Best for reasoning, coding, agent scenarios. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Perf: ~96.8 tok/s · first token ~1.4s
Best for reasoning, coding, agents. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~20 GB
Best for: Reasoning, Coding, Agent scenarios·Perf: ~96.8 tok/s · first token ~1.4s
Best for reasoning, coding, agent scenarios. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 27B / Q4_K_M / ~16.5 GB
Best for: Coding, Agent, Vision, Long context·Perf: ~36.8 tok/s · first token ~0.7s
Best for coding, agent, vision, long context. Strong fit for 128 GB RAM with balanced speed and quality.
Qwen / 27B / Q4_K_M / ~18 GB
Best for: Coding, Quality, Long context·Perf: ~36.8 tok/s · first token ~0.7s
Best for coding, quality, long context. Strong fit for 128 GB RAM with balanced speed and quality.
The wizard tunes picks and speed estimates to your exact device, chip, and RAM.