Llama 4 Maverick
Llama / 400B / Q4_K_M / ~245 GB
Best for: Frontier quality, Long context·Perf: ~8 tok/s · first token ~2.5s
Best for frontier quality, long context. Strong fit for 512 GB RAM with balanced speed and quality.
Ranked open-weight models that run well on a 512GB machine, with estimated speed and the exact ollama command for each.
Try: RTX 4090, MacBook Pro M4, 16GB, iPhone 17 Pro
Estimates assume a representative Apple-Silicon machine with 512GB unified memory. Tok/s are ModelFit estimates, not measured benchmarks. Run the wizard for figures tuned to your exact chip.
Llama / 400B / Q4_K_M / ~245 GB
Best for: Frontier quality, Long context·Perf: ~8 tok/s · first token ~2.5s
Best for frontier quality, long context. Strong fit for 512 GB RAM with balanced speed and quality.
GPT-OSS / 117B / MXFP4 / ~65.4 GB
Best for: Reasoning, Coding, Agents·Perf: ~28.6 tok/s · first token ~1.6s
Best for reasoning, coding, agents. Strong fit for 512 GB RAM with balanced speed and quality.
Laguna / 118B / Q4_K_M / ~96 GB
Best for: Agentic coding, Long-horizon tasks·Perf: ~21.5 tok/s · first token ~1.7s
Best for agentic coding, long-horizon tasks. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 122B / Q4_K_M / ~72 GB
Best for: Frontier-level reasoning, Complex tasks·Perf: ~18.9 tok/s · first token ~1.8s
Best for frontier-level reasoning, complex tasks. Strong fit for 512 GB RAM with balanced speed and quality.
Llama / 109B / Q4_K_M / ~67 GB
Best for: Long context, Quality, Multimodal·Perf: ~15.4 tok/s · first token ~1.9s
Best for long context, quality, multimodal. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Perf: ~64.6 tok/s · first token ~1.4s
Best for reasoning, coding, agents. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~20 GB
Best for: Reasoning, Coding, Agent scenarios·Perf: ~64.6 tok/s · first token ~1.4s
Best for reasoning, coding, agent scenarios. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 235B / Q4_K_M / ~130 GB
Best for: Quality, Reasoning·Perf: ~9.2 tok/s · first token ~2.3s
Best for quality, reasoning. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 80B / Q8_0 / ~84.8 GB
Best for: Chat, Coding, Long Context·Perf: ~23.4 tok/s · first token ~1.7s
Best for chat, coding, long context. Strong fit for 512 GB RAM with balanced speed and quality.
Gemma / 26B / Q4_K_M / ~16 GB
Best for: Chat, Coding, Multimodal·Perf: ~64.9 tok/s · first token ~0.6s
Best for chat, coding, multimodal. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 4B / Q4_K_M / ~3.5 GB
Best for: Coding, Agents, Multimodal·Perf: ~165.4 tok/s · first token ~0.5s
Best for coding, agents, multimodal. Strong fit for 512 GB RAM with balanced speed and quality.
Qwen / 27B / Q4_K_M / ~16 GB
Best for: Chat, Coding, Complex reasoning·Perf: ~24.5 tok/s · first token ~0.9s
Best for chat, coding, complex reasoning. Strong fit for 512 GB RAM with balanced speed and quality.
The wizard tunes picks and speed estimates to your exact device, chip, and RAM.