By Peter · ModelFit · 2026-09-11

Llama 4 GPU Requirements: Scout and Maverick (2026)

A workstation panel with an illuminated Meta infinity-loop sign and a tall wall of memory modules, representing Llama 4 needing 96GB+

Llama 4 has the largest memory requirements of any widely used open family, and the gap between the two variants is the story. Scout (109B total, 17B active) needs 67 GB at Q4. Maverick (400B total, 17B active) needs 245 GB. Neither fits a consumer GPU, and we verified both figures against the ModelFit engine on September 3, 2026 before publishing this. If a sizing guide shows either on a 24GB card, it is counting CPU offload as running.

TL;DR: Scout loads in 67 GB, so the floor is a 96GB machine: Mac Studio M5 Max/Ultra or RTX PRO 6000. Maverick loads in 245 GB, so it wants a 256GB Mac Studio M5 Ultra and nothing smaller. There is no consumer GPU path to Maverick. Both are MoE with 17B active, so once loaded they decode faster than their size suggests.

How much VRAM do Scout and Maverick need?

VariantArchitectureLoaded (Q4_K_M)Min machine10M context?
Llama 4 Scout109B MoE, 17B active~67 GB96 GB unified / VRAMYes, on 96GB+
Llama 4 Maverick400B MoE, 17B active~245 GB256 GB Mac StudioNo (1M context)

Scout's headline feature, per Meta's Llama site and the model cards on meta-llama on Hugging Face, is its 10M-token context window. That context is not free: even with efficient caching, long-document work on Scout wants the full 96GB budget. Maverick tops out at a 1M context and is a 256GB-machine model, full stop. Both variants have their own model pages here: Scout and Maverick.

Which hardware actually runs Llama 4?

ModelFit engine estimates, labeled est.

HardwareMemoryScout est. tok/sMaverick
Mac Studio M5 Ultra 512GB512 GB~23 est.Fits, slow
Mac Studio M5 Ultra 256GB256 GB~23 est.Fits at 245 GB, no headroom
Mac Studio M5 Ultra 96GB96 GB~23 est.No
Mac Studio M5 Max 96GB96 GB~17 est.No
RTX PRO 6000 96GB96 GB~35 est.No
RTX 5090 32GB32 GB~12 est. (offload)No
RTX 4090 24GB24 GB~9 est. (offload)No

Consider the 4090 and 5090 Scout rows: those are offloaded numbers, with two-thirds of the model sitting in system RAM. Some sites grade this configuration highly. We ran the same configurations through the engine and, at ~9-12 tok/s est. with heavy latency, we do not recommend them. A full ranked list for Scout lives on its best-hardware page.

Why is the Mac Studio the default Llama 4 machine?

Both Llama 4 variants are unified-memory plays. A Mac Studio M5 Ultra offers 96GB, 256GB, or 512GB of memory that the GPU addresses directly at 1.2 TB/s. No consumer GPU offers that pool, and a 96GB workstation card (RTX PRO 6000) costs more than the Mac Studio that matches it. For Scout, the choice between a 96GB Mac Studio and an RTX PRO 6000 comes down to ecosystem; both run it well. For Maverick, the 256GB Mac Studio is effectively the only option short of a multi-GPU server.

What should you run instead on normal hardware?

The honest answer for 24-48GB hardware is that Llama 4 is not the pick. The models that deliver comparable everyday quality within reach:

  • Qwen 3.5 122B-A10B (~72 GB) — the frontier MoE for 96GB machines, slightly heavier than Scout, competitive quality. See the 122B requirements guide.
  • GPT-OSS 120B (~65 GB) — the lightest 120B-class option, fits 96GB with more headroom than Scout. See the GPT-OSS 120B model page.
  • Qwen 3.6 35B-A3B (~22 GB) — the 32GB-tier MoE that captures much of the agentic quality at a third of Scout's memory.
  • GPT-OSS 20B (~14 GB) — for 24GB machines, the best quality per GB in this class. See the GPT-OSS 20B model page.

Why do we not list a consumer GPU path at all? Because there is none, and pretending otherwise is how people end up with an expensive card that loads Llama 4 instead of running it.

FAQ

Can a 24GB or 32GB GPU run Llama 4 Scout?

Not practically. Scout is 67 GB. On a 24GB card it offloads most of its weights to system RAM and runs at roughly 9-12 tok/s est. with heavy stalls. It loads; it does not run well.

What runs Llama 4 Maverick?

A Mac Studio M5 Ultra at 256GB or 512GB, or a multi-GPU server. There is no single consumer or workstation GPU with 245 GB of VRAM. Maverick is the one mainstream open model that is effectively Mac-only at the high end.

Is Scout's 10M context real?

Yes, and it is the main reason to choose Scout over the lighter 120B alternatives. It requires the full 96GB budget to use seriously. On smaller memory the model fits but the context that makes it special does not.

Scout or GPT-OSS 120B?

GPT-OSS 120B is lighter (about 65 GB vs 67 GB) and slightly faster on the same hardware. Scout wins if you need the 10M context window or Meta's multimodal tuning. For pure reasoning per GB, GPT-OSS 120B is the more efficient pick.

___

Every figure in this guide comes from the ModelFit recommendation engine, re-run on September 3, 2026, with speeds labeled est. How we work: about ModelFit. Corrections welcome: contact.
What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

The weekly local-AI refresh

New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.

By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.

Have questions? Reach out on X/Twitter