A 16GB Mac is the most common local AI machine, and the most misconfigured. macOS reserves a share of memory before any model loads, so the usable pool is about 11 GB, not 16. That maps cleanly to 8B-12B models at Q4_K_M, and the 2026 model generation made this tier genuinely good. We tested and ranked these picks with the ModelFit engine on September 3, 2026, sized for the real 11 GB budget — not the number on the box.
TL;DR: On 16GB, Qwen 3.5 9B is the top pick (~25 tok/s est. on an M6 Mac Mini). Qwen 3 8B and Llama 3.1 8B are the fast proven options at ~28 tok/s est. Gemma 4 12B is the largest comfortable fit at ~19 tok/s est. Ornith 1.0 9B is the agentic coding pick. Skip 14B: it loads but leaves no room for context.
What is the real memory budget on 16GB?
macOS does not let you use all 16 GB for a model. In fact, the GPU-wired memory cap and the OS together take the first slice before your runner even starts.
| Allocation | Typical size |
|---|---|
| macOS kernel + services | ~2-3 GB |
| Active apps (browser, terminal) | ~1-3 GB |
| Available for LLM | ~11 GB |
With Q4_K_M costing roughly 0.6 GB per billion parameters, 11 GB fits a 12B model tightly or a 9B model comfortably with context. A 14B, however, needs about 11 GB for the weights alone, which leaves nothing for a single token of context. It loads and then swaps. Why do so many guides still recommend 14B on 16GB? Because loading looks like running right up until the swap storm starts.
Bandwidth sets the speed once the model fits. An M5 or M6 at 153 GB/s runs the 7B reference at roughly 32 tok/s est., while an older M4 at 120 GB/s runs the same model at roughly 24 tok/s est. That said, even the slower tier stays perfectly usable for chat.
Which models rank best on 16GB?
Speeds are ModelFit engine estimates on a Mac Mini M6 16GB (153 GB/s), labeled est. A MacBook Air M5 16GB runs about 10% slower on the same models because it is fanless.
| Rank | Model | Size | Est. tok/s | Best for |
|---|---|---|---|---|
| 1 | Qwen 3.5 9B | ~7 GB | ~25 est. | Quality, coding, reasoning |
| 2 | Qwen 3 8B | ~6.5 GB | ~28 est. | Chat, coding |
| 3 | Gemma 4 12B | ~8 GB | ~19 est. | Multimodal, quality |
| 4 | Ornith 1.0 9B | ~5.6 GB | ~25 est. | Agentic coding |
| 5 | Llama 3.1 8B | ~6.5 GB | ~28 est. | Ecosystem default |
| 6 | LFM2.5 8B-A1B | ~5.5 GB | ~64 est. | Agents, tool calling |
What is each pick for?
Qwen 3.5 9B is the best single answer. It rivals 30B models from a year ago on reasoning, handles text and images, and still leaves 4 GB of headroom for context on a 16GB machine. For example, a full coding session with a few thousand lines of context stays entirely in memory. The tracked build matches Qwen's Hugging Face org; see the Qwen 3.5 9B model page.
Qwen 3 8B is the fastest proven option and the one every tool supports, from Ollama to LM Studio to raw llama.cpp. At ~28 tok/s est. it is the responsive daily driver, and because it has been around long enough for every runtime to optimize it, it is also the least likely pick to surprise you. See the Qwen 3 8B model page.
Gemma 4 12B is the quality ceiling at this memory, and the tradeoff is explicit: it is the largest dense model that fits cleanly, bought at the cost of speed — ~19 tok/s est. on an M6, ~17 on an M5 Air. If your sessions are long documents rather than quick chats, that tradeoff is usually worth making. See the Gemma 4 12B model page.
Ornith 1.0 9B loads in just 5.6 GB, the smallest footprint here, leaving the most room for long agent contexts and tool outputs. See the Ornith model page.
What about the fanless MacBook Air specifically?
Everything in the ranked table applies to the Air, with one thermal footnote that matters more than the spec sheet suggests. The Air has no fan, so a long agent session or a big batch of document summaries will pull its sustained speed down by about 10% compared to the Mac Mini on the same chip, while short interactive prompts are unaffected. In practice that means the Air is the right machine for chat, writing, and occasional coding help, and the Mini is the right one if an agent will grind for hours.
Two habits recover most of the difference without spending anything: keep contexts moderate rather than maxed, and close the browser tabs you are not using, because every gigabyte of RAM the system reclaims is a gigabyte the model can use for context. Neither changes the picks; both change how the picks feel.
What can 16GB not run?
- 14B dense at Q4 (Qwen 3 14B, Phi-4): ~11 GB of weights leaves no context headroom. Loads, then swaps under real use.
- 20B-class MoE (GPT-OSS 20B at about 14 GB): over budget. This is the cleanest argument for the 24GB config. We tested both memory tiers side by side, and the 20B class is where the experience gap between 16GB and 24GB is largest.
- 27B and up: no.
- High quants: even the 9B class at Q8 (about 11 GB) is tight on 16GB.
If you need 14B+ or GPT-OSS 20B comfortably, the 24GB M6 Mac Mini is the meaningful step up; the machine itself is on Apple's Mac Mini specs page. For the full model table at 16GB, see also the 16GB Mac Mini M6 breakdown.
FAQ
Is 16GB enough for local AI in 2026?
Yes, for 8B-12B models — the 2026 small-model generation (Qwen 3.5 9B, Gemma 4 12B) is good enough that 16GB covers most chat, coding, and assistant work. It is not enough for 14B+ or the 20B-class MoE models, and that ceiling is worth knowing before you buy.
MacBook Air or Mac Mini at 16GB?
Same models, different sustained speed: the fanless Air throttles about 10% under long sessions, while the Mini holds its speed indefinitely. For occasional use both are fine; for all-day agent work, the Mini.
Should I upgrade from 16GB to 24GB?
If you want GPT-OSS 20B, 14B dense models, or long-context agent workflows, yes — 24GB is the tier where the 20B-class MoE models fit, and they are the biggest quality-per-GB story of 2026. If you are happy with 8B-9B chat, 16GB is enough.
What is the fastest model on 16GB?
LFM2.5 8B-A1B at ~64 tok/s est. on an M6, because its 1.5B active path reads few bytes per token. Among dense models, Qwen 3 8B and Llama 3.1 8B tie at ~28 tok/s est.
___
Rankings and speeds in this guide come from the ModelFit engine, re-run on September 3, 2026, with estimates labeled est. How we work: about ModelFit. Corrections welcome: contact.Where to Buy for Local AI
best configsPrefer to buy direct? Buy from Apple (same price, no affiliate link).
Archive your model library off the internal drive. Quantized models run 5 to 40GB each, so 2TB holds dozens with room to spare.
Check price on Amazon40Gbps external storage fast enough to run models from. Pair it with an M.2 drive for a portable model vault.
Check price on AmazonThe fanless MacBook Air heat-soaks on long inference runs. An aluminum riser lifts the chassis so it sheds heat better off the desk.
Check price on AmazonMore ports for the external drives, displays and peripherals around a local-AI workstation.
Check price on AmazonModelFit may earn a commission on purchases through these links, at no extra cost to you.
Want a Model Bigger Than This Mac Runs? Rent a Cloud GPU
by the hour70B+ and frontier open-weight models that won't fit in unified memory run great on an hourly rented GPU, same open weights, same Ollama workflow, no subscription.
ModelFit may earn a commission on sign-ups made through these links, at no extra cost to you.
The weekly local-AI refresh
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Have questions? Reach out on X/Twitter