Ollama v0.40.0 changes which runtime your Mac uses, not which models exist. On Apple Silicon, model architectures that the MLX runtime supports now run on MLX automatically, so a model that used to land on the older llama.cpp path can now land on Apple's own array framework with no flag to set. The release note names qwen3.8, gemma4, qwen3.6 and qwen3.5, plus four decision models and one embedding model. The question for anyone whose Mac already runs Ollama is narrower: does this update change what happens on your machine, and how do you confirm it? This article separates the routing change from the speed claim most people will assume comes with it.
TL;DR: Ollama v0.40.0 is a stable release, not a prerelease, and it switches Apple Silicon to MLX by default for model architectures the MLX runtime supports. The note names qwen3.8, gemma4, qwen3.6 and qwen3.5, moves the decision models nimble, tev1, clef and clef-flash to MLX, and adds embeddinggemma-2 as the first embedding model on MLX. There is no benchmark and no tokens per second figure in the note, so Ollama claims no speedup here. Models on architectures MLX does not support yet keep running on the previous path.
What Ollama v0.40.0 actually changed
MLX is Apple's array framework, and the runtime Ollama added for Apple Silicon alongside llama.cpp. Until this release, an Apple Silicon user reached MLX only when the model shipped an MLX-tagged variant on the library or when the model already routed there. As our write-up of v0.34.4 put it, MLX was still not the default runner at that point, and the v0.40.0 release candidate was the build that would flip it. v0.40.0 is that flip.
The release facts first, because they decide whether to update:
- GitHub records v0.40.0 as a stable release. The release record carries
draft: falseandprerelease: false. - It is the newest tag in the repository, ahead of v0.35.1. The tag list also holds the v0.40.0-rc0 through v0.40.0-rc6 candidates, so the stable build is the end of that line, not the start.
- The release page carries a 2026-09-25 publication stamp and a 2026-10-06 record update. Treat the changelog, not the stamps, as the source of truth for what is inside.
- The changelog compares v0.35.1 to v0.40.0 rather than to v0.34.4, and that diff is 26 commits across 206 files.
None of that touches Linux or Windows. On those platforms the MLX runtime is not the path Ollama builds on, so the default switch is an Apple Silicon story.
Which model architectures now route to MLX by default
The note defines the switch at the architecture level: on Apple Silicon, models whose architecture the MLX runtime supports run on MLX automatically. Ollama names these as the models you can pull today and expect on MLX:
- qwen3.8, the model the note uses as its example, at 27B with vision and a 256K context window.
- gemma4, across its e2b, e4b, 12b, 26b and 31b sizes.
- qwen3.6 at 27B and 35B.
- qwen3.5, from 0.8b to 122b.
- The decision models nimble, tev1, clef and clef-flash.
- embeddinggemma-2, the first embedding model Ollama runs on MLX.
The list is not exhaustive, and Ollama does not publish the full set of architectures the runtime supports. The note closes with "We will continue testing and enabling additional models", which means the boundary moves again in later releases. If your model is not on the list and its architecture is not supported by the MLX runtime yet, it does not move, and it keeps running exactly as it did before.
The practical consequence: the routing is keyed to the architecture, so tags of a supported family should inherit it rather than needing one by one enrollment. A qwen3.5 tag you already have on disk should change runner after the update, without a re-pull of the model weights.
Get told when a better model fits your MacBook Air M5 16 GB
One email when a new open-weight model is a better fit for your machine, plus the Thursday weekly.
For: MacBook Air M5 16 GB
Does updating make your Mac faster? What Ollama claims, and what it does not
This is where the release note is silent, and where most coverage will fill the gap with a number.
- The note contains no benchmark, no tokens per second figure and no before and after comparison. It states a routing change and nothing more.
- ModelFit runs no benchmarks of its own. We did not produce a v0.34.4 versus v0.40.0 measurement for this article, so no measured speedup appears here either.
- The honest expectation is that the win equals whatever the MLX path already gave you for that architecture. If you were running the MLX-tagged variant of a model before, or the model was already routed to MLX, this release changes nothing for you. If you were on the llama.cpp path for a supported architecture, you move onto MLX, and MLX on Apple Silicon has generally been the faster of the two paths. That is a direction, not a number.
Anyone quoting a percentage gain for v0.40.0 on Apple Silicon is quoting something outside the release. Ask for the method and the machine before you believe it.
The explicit MLX tag is no longer the only way in
Before this release, the way to pin a model to MLX was to pull the -mlx tag where the library published one. Those tags still exist, and the newly supported families are the same ones that carry them:
| Family | Explicit MLX tag | Size on the library page |
|---|---|---|
| qwen3.8 | qwen3.8:27b-mlx | 18GB |
| qwen3.6 | qwen3.6:27b-mlx, qwen3.6:35b-mlx | 19GB, 24GB |
| qwen3.5 | qwen3.5:9b-mlx, qwen3.5:27b-mlx, qwen3.5:35b-mlx | 8.9GB, 20GB, 22GB |
If you want the runner to be explicit and not depend on the default routing rule, the -mlx tag remains the way to force it. If you just want the faster path, the default now gets you there for the supported architectures.
Can your Mac run the newly MLX-routed models?
The routing change is about the runner, not about memory. A model that did not fit your Mac before v0.40.0 does not fit after it. The flip side is that if one of these models was already on your shortlist, the update removes the tag choice from the setup.
The budget below comes from the ModelFit engine, which computes the minimum unified memory a model needs on Apple Silicon. Sizes and context windows are on each library page linked above.
| Mac unified memory | Models from the v0.40.0 list that fit | Engine minimum RAM |
|---|---|---|
| 16GB | gemma4:12b, gemma4:e4b, qwen3.5:9b | 12GB, 6GB, 12GB |
| 24GB | qwen3.8:27b, gemma4:26b, qwen3.5:27b | 24GB |
| 32GB | qwen3.6:27b, qwen3.6:35b-a3b, qwen3.5:35b-a3b, gemma4:31b | 32GB |
| 48GB | the Q8 builds of the 27B and 31B models above | 48GB |
| 96GB and up | qwen3.5:122b-a10b | 96GB |
Two caveats on that table. First, the engine budget is the floor for the model plus its context, and a busy Mac needs headroom above it, so a 24GB machine running qwen3.8:27b is at the line, not comfortably above it. Second, the memory number does not change with the runner in any documented way. Ollama does not publish an MLX-versus-llama.cpp memory delta for these architectures, so treat any claim that MLX saves or costs you memory on a Mac as unverified.
embeddinggemma-2 on MLX: what it enables for local RAG
The embedding line in the note is the smallest item and the one most likely to matter for a local pipeline. embeddinggemma-2 is Google DeepMind's multimodal embedding model built on the Gemma 4 architecture. It maps text, code, images, video and audio into a single 768-dimensional space and supports Matryoshka truncation at 128d, 256d, 512d and 768d.
On a Mac, that combination does three things:
- Retrieval augmented generation stops needing a cloud embedding call. You index documents locally with
ollama pull embeddinggemma-2and query the same Mac. - The footprint is small enough to sit beside the generator. The library lists the 270m tag at 378MB, the 440m at 714MB, the 570m at 990MB and the 740m at 1.3GB.
- Running the embedding model on MLX puts the indexing step on the same runtime as the generation step for the supported chat models, which keeps one memory model in play instead of two.
If you were choosing between a terminal-based Ollama stack and a desktop app for this, the app route is documented in our LM Studio guide for Mac, but the embedding model above is an Ollama library artifact and runs through Ollama first.
How to update, and how to confirm the runner
- Update the CLI with
brew upgrade ollama, or take the installer from ollama.com/download. Our Ollama install guide for Mac covers the clean first install. - Check you are actually on the new line with
ollama --versionand confirm it reads 0.40.0 rather than a release candidate. A desktop app update and abrew upgradecan sit at different versions, so check both if you use the app. - Restart the server after the update. The running
ollama serveprocess keeps the old binary until it is restarted, and so does the app. - To confirm what loaded, pull a model from the list and watch the server log at
~/.ollama/logs/server.log, the path Ollama documents for macOS. If you want the choice to be explicit rather than inferred, pull the-mlxtag.
Ollama does not document a command that prints the active runner for a loaded model, and there is no environment variable in the published docs that forces MLX across the board. Not verified means not to be trusted: if a guide tells you to set a flag to turn MLX on globally, it is describing behavior the docs do not state.
Who sees no change at all
Four groups get nothing measurable from this release:
- Anyone on Intel Macs, Linux or Windows, because the switch is Apple Silicon only.
- Anyone whose models sit on architectures the MLX runtime does not yet support, which Ollama does not enumerate.
- Anyone already running the MLX-tagged variant of a supported model, since they were on that path before the update.
- Anyone working through the cloud path, or through a desktop app that pins its own runtime version. The MLX comparison across runtimes is covered in MLX versus Ollama on Mac, and the decision-model angle in our Laya piece.
If you fall into one of those groups, v0.40.0 is a safe but uneventful update. If you run a supported Qwen or Gemma model on an Apple Silicon Mac on the llama.cpp path, it is the release that moves you.
FAQ
Is Ollama v0.40.0 a stable release or a prerelease?Stable. The GitHub release record for v0.40.0 carries draft: false and prerelease: false, and it is the newest tag in the repository ahead of v0.35.1. The v0.40.0-rc0 through v0.40.0-rc6 tags are the candidates that came before it. The release page carries a 2026-09-25 publication stamp and a 2026-10-06 record update.
No. The switch applies to model architectures the MLX runtime supports on Apple Silicon. Ollama names qwen3.8, gemma4, qwen3.6, qwen3.5, the decision models nimble, tev1, clef and clef-flash, and the embedding model embeddinggemma-2. Everything else keeps running on the previous path until Ollama enables it.
Does Ollama claim a speedup in v0.40.0?No. The release note contains no benchmark and no tokens per second figure. It describes a routing change. ModelFit runs no benchmarks, so no measured before and after number appears in this article either.
Do I still need the-mlx tag?
Only if you want the runner to be explicit. For the supported architectures, MLX is now the default, so pulling the plain tag gets the same runtime. The -mlx tags remain published for the families that carry them.
An Apple Silicon Mac running a supported Qwen or Gemma model that was on the llama.cpp path. On a 16GB machine that means gemma4:12b or qwen3.5:9b, on a 32GB machine the 27B and 30B-class builds, and on a 96GB machine up to qwen3.5:122b-a10b. Macs already on MLX, and models on unsupported architectures, see no change.
Sources
Don't miss the next one for your MacBook Air M5 16 GB
This article covers one moment. Get one email when a newer open-weight model fits your machine better, plus the Thursday weekly with what changed.
For: MacBook Air M5 16 GB
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
Have questions? Reach out on X/Twitter