Cloudflare uploaded Clef and Clef Flash to Hugging Face on 30 September 2026. For Mac users, therefore, the answer splits cleanly by memory tier. Clef Flash, the 9B model, runs on a 16 GB Apple Silicon Mac through a community 4-bit MLX build. The 27B Clef needs a 32 GB machine at 4-bit and 48 GB at 8-bit. Cloudflare publishes no Apple Silicon figure of any kind, however. We found none in the release notes or in either Hugging Face card. In contrast, the GGUF quants that appeared within a day cannot score options at all yet. The joint schema head, the piece that turns hidden states into probabilities, never made it into a GGUF file, and llama.cpp support for the model, consequently, is still an open pull request. So what follows is what a decision model actually returns, which Mac takes which build, and whose latency numbers are whose.
TL;DR: Clef Flash 9B at 4-bit MLX is a 6.2 GB download that peaks at 7.0 to 8.6 GB, and the MLX card documents a 16 GB Mac. Clef 27B at 4-bit MLX is a 16.3 GB download that peaks at 17.1 to 19.6 GB, and the card documents a 32 GB Mac, so it does not fit 16 GB or 24 GB. Ollama serves both through /v1/systemone from v0.35.1 (29 September 2026); clef is an 18 GB pull and clef-flash is 11 GB. The bartowski and ggml-org GGUF quants carry the backbone and the vision projector but no joint head, and llama.cpp's Clef support (PR 29831) is unmerged, so LM Studio and plain llama.cpp generate text instead of scoring options.
What Clef returns, and what it never does
A decision model is not a chat model. Clef reads a state (text, JSON, images or video) plus a schema of typed questions, and returns one probability per allowed option in a single forward pass. There is no generated text at all, and there is nothing to extract from a completion. Both checkpoints run three question types:
choice: named options, returned with a probability map over the options plus a confidence value.score: an ordered rubric, returned as a probability-weighted number plus its legend.noul: the probability that a condition is true.
The release ships a working request example in joint_schema_model.py. One call carries all three types at once, over the same state.
from joint_schema_model import systemone
response = systemone(model, processor, {
"model": "clef-flash",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
The answer body is the part readers expect to be text, and it is not. Ollama's release notes show the shape for a single noul question. The response carries the question id, its type, and the probability (0.958), with output_tokens at 0. For a choice question, the same body carries the winning option, a per-option probability map and a confidence value. For a score question, it is a weighted number between the rubric levels. In other words, nothing in the response is free-form prose. As a result, there is no parsing step and no reasoning wait.
If you were expecting a chat model, this is the whole surprise. Ollama even marks the capability of these artifacts as decision only, so its clients stop offering them for chat, tools or thinking.
Does Clef run on your Mac?
The short answer, by unified memory tier:
| Mac unified memory | Clef Flash 9B, 4-bit MLX | Clef Flash 9B, 8-bit MLX | Clef 27B, 4-bit MLX | Clef 27B, 8-bit MLX |
|---|---|---|---|---|
| 16 GB | Runs, documented minimum | Short prompts only | Does not fit | Does not fit |
| 24 GB | Runs | Runs, documented minimum | Does not fit | Does not fit |
| 32 GB | Runs | Runs | Runs, documented minimum | Does not fit |
| 48 GB and up | Runs | Runs | Runs | Runs, documented minimum |
Get told when a better model fits your MacBook Air M5 16 GB
New open-weight models land every week. When one beats today's pick on your exact machine, you get one short email: the model, why it is better, and the command to run it.
You'll also get the weekly ModelFit email. By signing up you agree to our Privacy Policy. Unsubscribe from either email anytime, in one click.
Unified memory is the single pool of RAM that Apple Silicon shares between the CPU and the GPU, which is why the ceiling for a model is a RAM number rather than a VRAM number. For instance, a 24 GB Mac runs the 8-bit Flash build but cannot take the 27B at 4-bit. In particular, every minimum in that table, and every Apple Silicon memory and latency figure in this article, comes from the community MLX conversion, which documents its own table. Here is the same source unrolled by build:
| Build | Base | Download | Peak memory, 1k to 16k tokens | Latency, 1k / 16k tokens | Minimum Mac RAM |
|---|---|---|---|---|---|
| clef-flash-4bit | 9B | 6.2 GB | 7.0 to 8.6 GB | 0.31 s / 7.0 s | 16 GB |
| clef-flash-8bit | 9B | 10.7 GB | 11.4 to 13.0 GB | 0.34 s / 7.7 s | 24 GB (16 GB for short prompts) |
| clef-4bit | 27B | 16.3 GB | 17.1 to 19.6 GB | 1.4 s / 26.0 s | 32 GB |
| clef-8bit | 27B | 29.8 GB | 30.5 to 33.0 GB | 1.5 s / 32.0 s | 48 GB |
Two important caveats apply to that table. The author measured the figures on an M5 Max with text input, so image and video inputs cost more than the numbers suggest. In addition, macOS lets the GPU use only about 70 to 75 percent of RAM by default, and that is why the documented minimum sits well above the peak.
None of it is Cloudflare's or ours. Their authors tested the Hugging Face Cloudflare cards on a single H200, and Cloudflare tested on its own GPUs. Neither source publishes an Apple Silicon reading. This site publishes no benchmarks; consequently, the only Mac numbers on this page are the MLX build author's, quoted with that label. In our experience, that gap between peak memory and the documented minimum is what catches people out when they buy a 16 GB Mac for a 27B model.
In fact, the unquantised checkpoints are not laptop material at all. The Clef repository totals 54,989,894,057 bytes, and Clef Flash totals 19,083,377,402 bytes, both BF16. That is precisely why every local path goes through a quantised build. For the memory arithmetic behind those tiers, see how much VRAM local LLMs need. We built the same budget check into our recommender, which scores weights plus KV cache against the share of unified memory that macOS leaves to the GPU.
The GGUF question: a quant that cannot score options
This is the part that decides whether the local story is real or nominal. Specifically, Clef's typed output comes from a joint schema head, a small transformer that reads the backbone's final hidden states and scores all options of all questions together. A joint schema head is a separate file, joint_head.safetensors, 256,125,024 bytes for Clef and 243,538,016 bytes for Clef Flash.
The GGUF quants, however, do not include it. The bartowski conversions, published on 1 October 2026 and already past three thousand downloads for Flash, ship quantised weights plus the multimodal projector files. They carry no joint head at all. Notably, their model card documents a conventional chat and tool-calling prompt format, not a /v1/systemone request. The ggml-org conversions say the quiet part out loud in their own cards: "This is a decision model, to be used via /v1/systemone API. Requires [llama.cpp PR 29831]". That pull request, which adds text-only Clef support, remains open and unmerged. PR 29832, the model-agnostic rewrite of the server's decision API, remains open too. In contrast, the /v1/systemone path that did merge on 2 October covers laya, julia-1, lev, openjev and kev, and it does not cover Clef.
Conversely, the MLX conversion reaches the opposite conclusion by shipping the missing piece. Its repositories carry joint_head.safetensors and the bundled clef_mlx.py loader. Its model card is blunt about the alternative: "It is not a chat model. mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the joint schema head."
So the honest local picture today is two paths, not four:
- Ollama from v0.35.1, through
/v1/systemone, which is the documented, vendor-supported route. - The MLX build's
clef_mlx.py, which exposespredict,systemoneand an HTTP server on/v1/systemone.
Plain llama.cpp and LM Studio, therefore, are not there yet. A GGUF quant is a single-file build that carries the backbone and the vision projector only. It will load and generate text. The text it generates means nothing, because the scorer is not present in the file. For example, the bartowski Flash quants open fine in LM Studio and still answer in prose. Meanwhile, the ggml-org GGUF repos for both models show zero downloads, and so does the third-party Livesport GGUF conversion. As a result, nobody has validated a local GGUF path in practice either.
Cloudflare's 209.3 ms is not a Mac number
Cloudflare's own latency table is genuine, and it is not about your laptop. Across the 43 benchmark evaluations it performed, it reports a median latency of 209.3 ms for Clef, 38.8 ms for Clef Flash and 524.1 ms for Jev. The p95 figures are 238.6 ms, 122.4 ms and 536.0 ms. Cloudflare measured those on its own serving stack, on Workers AI GPUs behind its edge. That is exactly what the following paragraph about hosting on Workers AI is advertising. Notably, there is no Apple Silicon figure anywhere in the release.
The Mac figures worth having are the MLX build's, and they are, in fact, much smaller at short prompts. Clef Flash 4-bit returns a decision in 0.31 s at 1,000 tokens and 7.0 s at 16,000 tokens. The 27B 4-bit takes 1.4 s and 26.0 s for the same two lengths. Read those as one machine's spot check on a single M5 Max, not as a floor or a ceiling.
For the rest of the decision-model field, our earlier write-ups cover four angles. One is what Ollama's decision lane costs on a Mac. Another is why a 421M decision encoder can answer without a GPU. Similarly, we covered the open Kev family and its 96 GB ask and the prior art behind Jev's closed API.
Licence, base model, and what else you are downloading
Notably, both checkpoints are Apache-2.0, and both are post-trains of Qwen backbones, so the download is not a from-scratch Cloudflare model:
- Clef is a 27B post-train of Qwen3.8-27B. The repository holds 25 files, twelve weight shards, the joint head, the tokenizer, the image and video processor, and the custom code in
joint_schema_model.py. - Clef Flash is a 9B post-train of Qwen3.5-9B, in four weight shards plus the same supporting files.
- Ollama serves both:
clefis an 18 GB pull andclef-flashis 11 GB, and the library page states that Clef requires Ollama 0.35.1 or later. - Adoption at the time of writing: 824 downloads and 768 likes for Clef, 1,303 downloads and 256 likes for Clef Flash. On the MLX side, clef-flash-4bit has 215 downloads and 12 likes, clef-4bit has 110 downloads and 4 likes, and both 8-bit builds are under ten downloads, which is the usual pattern for the first week of a release.
One sizing warning before you download anything. Ollama's variant table lists a 256K context window, while the same page's highlights describe a 64K window. In addition, the release code bounds a request at 16,384 tokens by default through encode_record. Plan for the smaller number until you have verified the larger one.
FAQ
Is Clef a chat model?
No. It reads a state plus typed questions and returns probabilities, one per allowed option, in a single forward pass. The response carries no text; in fact, Ollama reports output_tokens at 0. If you want a chat model, you want the Qwen backbone underneath it, not Clef.
Can I run Clef in LM Studio or plain llama.cpp?
Not for decisions today. A GGUF quant loads the backbone and the vision projector, so LM Studio will generate plausible text from it. However, the joint schema head is not in the file and llama.cpp's Clef support is PR 29831, still open and unmerged. Use Ollama or the MLX build if you need typed probabilities.
Which Clef should I download for a 16 GB Mac?
Clef Flash at 4-bit MLX, a 6.2 GB download that peaks at 7.0 to 8.6 GB, and the MLX card documents a 16 GB machine. The 27B, however, does not fit at any quantisation on this page: its 4-bit MLX download alone is 16.3 GB, and the card documents a 32 GB Mac.
How big is the Ollama download and which version do I need?
clef is an 18 GB pull and clef-flash is 11 GB, and both need Ollama v0.35.1 or later. Ollama published that release on 29 September 2026, and it added them through /v1/systemone. The original BF16 checkpoints are larger still: 54,989,894,057 bytes for Clef and 19,083,377,402 bytes for Clef Flash.
Is Clef really open source?
Yes, both checkpoints are Apache-2.0, and both are post-trains of Apache-2.0 Qwen backbones. In addition, the licence covers the weights and the custom code that reads and scores the questions, which ships in the same repository as joint_schema_model.py.
Sources
- Cloudflare/clef on Hugging Face (model card, Decision Index table, single-H200 usage, joint head file, 16,384-token default bound).
- Cloudflare/clef-flash on Hugging Face.
- Cloudflare/clef repository file tree and sizes (54,989,894,057 bytes, twelve shards, 256,125,024-byte joint head).
- Cloudflare blog: Introducing Clef (43 benchmark evaluations, median and p95 latency table).
- Ollama library: clef (18 GB pull, 256K listing, requires Ollama 0.35.1, benchmark table).
- Ollama library: clef-flash (11 GB pull).
- Ollama v0.35.1 release notes (published 2026-09-29,
/v1/systemone, decision capability). - mlx-community/clef-flash-4bit (Mac memory and latency table, bundled
clef_mlx.py, "not a chat model" note). - mlx-community/clef-4bit (27B 4-bit sizes and minimum RAM).
- bartowski/Cloudflare_clef-flash-GGUF (quants and mmproj, no joint head).
- ggml-org/Clef-GGUF (needs
/v1/systemone, requires llama.cpp PR 29831). - Livesport/clef-flash-GGUF (zero downloads).
- llama.cpp PR 29831, Clef model support (open, unmerged).
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
Have questions? Reach out on X/Twitter