Apple announced the M5 Ultra Mac Studio on August 26, 2026, with pre-orders open today and shipping on September 22. The headline number for local AI is 512GB of unified memory at 1.2 TB/s of memory bandwidth, starting at $5,499. That is roughly 50% more bandwidth than the M3 Ultra, and it brings back the big-memory Mac after the 2026 DRAM-shortage cuts. Here is what that hardware actually runs, model by model, with speed estimates from the ModelFit engine.
TL;DR: The 512GB M5 Ultra is the first Mac that runs 120B-class MoE models and Llama 4 Maverick 400B entirely in memory at usable speed. In ModelFit estimates it delivers ~135 tok/s on a 7B Q4 reference, versus ~90 on the M3 Ultra. For 70B dense models with long context, nothing else under $10,000 comes close. If you already own a 256GB+ M3 Ultra, the upgrade math is weak. If you run large models on a 96GB machine or rent cloud GPUs weekly, this is the box.
What Apple announced
The Mac Studio line now splits into M5 Max and M5 Ultra. All figures below come from Apple's official spec page, published August 26, 2026.
| Config | Chip | Bandwidth | Memory | Price |
|---|---|---|---|---|
| Mac Studio M5 Max | 18-core CPU, up to 40-core GPU | 460-614 GB/s | 36-128GB | from $2,499 |
| Mac Studio M5 Ultra | 30-core CPU, 64-core GPU | 1.2 TB/s | 96GB | from $5,499 |
| Mac Studio M5 Ultra (max) | 36-core CPU, 80-core GPU | 1.2 TB/s | 256-512GB | up to $18,299 |
Two details matter for local LLM work. First, memory bandwidth: token generation is bandwidth-bound, so the jump from 819 GB/s (M3 Ultra) to 1.2 TB/s is the whole story on speed. Second, capacity: 512GB returns after Apple retired the 512GB and 256GB M3 Ultra tiers earlier in 2026.
Apple never publishes Neural Engine TOPS figures, and its "faster AI" claims reference older chips on ambiguous tests. Bandwidth and capacity are the honest numbers, so those are what we use.
What 512GB actually unlocks
Raw capacity changes which models fit at all. Weights are only part of the bill: KV-cache for context adds gigabytes on top, and it grows with every token of conversation.
| Model | Weights (quant) | Fits in 512GB? | Notes |
|---|---|---|---|
| Llama 3.3 70B | ~42GB (Q4) | Yes, huge headroom | ~13GB extra KV at 32k context |
| Qwen3.6 35B-A3B | ~22GB (Q4) | Yes | MoE, 3B active per token |
| gpt-oss-120b | ~65GB (MXFP4) | Yes | 120B-class MoE, fast decode |
| Llama 4 Maverick 400B | ~220GB (Q4) | Yes | Largest open-weight model that fits |
| DeepSeek-class 670B+ | ~370GB+ (Q4) | Tight | Possible at low quant, short context |
A 256GB config already runs 120B MoE models comfortably. The 512GB tier is what moves 400B-class MoE models from "cloud only" to "on my desk". That class of model scores near frontier proprietary systems on coding benchmarks, and it now runs offline.
Speed: what +50% bandwidth means in tok/s
These are ModelFit engine estimates (est.), derived from published bandwidth, not lab measurements. Real numbers will land when units ship on September 22.
| Workload | M3 Ultra (819 GB/s) | M5 Ultra (1.2 TB/s) |
|---|---|---|
| 8B Q4 (Qwen3 8B) | ~83 tok/s est. | ~124 tok/s est. |
| 14B Q4 (Qwen3 14B) | ~47 tok/s est. | ~71 tok/s est. |
| 70B Q4 dense (Llama 3.3 70B) | ~10 tok/s est. | ~14 tok/s est. |
| 120B MoE (gpt-oss-120b) | ~29 tok/s est. | ~43 tok/s est. |
| 400B MoE (Llama 4 Maverick) | not realistic | ~12 tok/s est. (512GB) |
The pattern to notice: MoE models benefit most. Decode speed tracks active parameters, not total size, so a 120B MoE with ~5B active params runs far faster than a dense 70B. On 1.2 TB/s hardware, gpt-oss-120b becomes a genuine daily-driver assistant rather than a demo.
Our engine's top pick for a 512GB M5 Ultra is Llama 4 Maverick at roughly 12 tok/s est. — the same model reads as not realistic on a 96GB config, which is exactly what the 512GB tier buys. That speed is slow for chat but remarkable for batch work: document analysis, code review over large repos, and overnight agent runs that never touch an API bill.
Should M3 Ultra owners upgrade?
Honestly, most should not. If your M3 Ultra has 256GB or 512GB, you already run the same model classes; you would be paying $5,499 for a speed bump of about 50%, not new capability.
The upgrade makes sense in two cases. You own a 96GB M3 Ultra and keep hitting the memory wall on 70B-plus models with real context lengths. Or you rent cloud GPU time weekly for large-model work, in which case the hardware pays for itself in months against typical A100 or H100 rental rates.
For everyone buying fresh, the calculus is simpler: the 96GB M5 Ultra at $5,499 is the value entry, and the 512GB tier is for people who already know why they need it.
FAQ
Can the M5 Ultra Mac Studio run a 70B model well?
Yes. A 70B Q4 model needs about 42GB for weights plus KV-cache for context, so even the 96GB base config runs it with large headroom. In ModelFit estimates, expect about 14 tok/s est., up from roughly 10 on the M3 Ultra. That is usable for interactive work and fine for batch jobs.
Is 512GB overkill for local AI?
For 7B-70B models, yes. The 512GB tier exists for 120B-class MoE models at long context, 400B-class models like Llama 4 Maverick, and multi-model workflows where you keep several large models loaded. If your work tops out at 70B, the 96GB or 256GB configs are the right size.
How does the M5 Ultra compare to an RTX 5090 for local LLMs?
They answer different questions. The RTX 5090 has 32GB of VRAM and higher raw bandwidth per dollar, so it runs 8B-32B models much faster per dollar spent. The M5 Ultra runs models that simply do not fit in 32GB. Buyers choosing between them should decide on model size first, then speed.
When can I buy the M5 Ultra Mac Studio?
Pre-orders opened August 26, 2026, and units ship starting September 22, 2026, per Apple's store. Base M5 Ultra pricing starts at $5,499 with 96GB; the 512GB configuration sits at the top of the range, up to $18,299 fully loaded.
Are the tok/s figures here measured?
No. Everything labeled est. comes from the ModelFit engine, which scales published memory bandwidth against a reference 7B workload. We will replace estimates with measured numbers once shipped units run our bench. The direction and rough size of the jump are safe; the exact decimals are not.
Verdict
The M5 Ultra Mac Studio is the most capable local-AI Mac Apple has sold, and the 512GB config reopens a model class that left the Mac lineup during the 2026 memory shortage. Buy the 96GB for 70B work, the 256GB for 120B MoE, and the 512GB if 400B-class models are your actual workload. Check the fit for your exact config on our Mac Studio M5 Ultra page before you pre-order.
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
The weekly local-AI refresh
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Have questions? Reach out on X/Twitter