CLM-8B is a decision model, not a chat model. A decision model is a network that reads one state plus a set of typed questions and returns labels or probability distributions, with no tokens generated. Contrastive-LM published this one on 21 September 2026 under Apache-2.0. The 4-bit encoder file in the community GGUF repo is 5,489,669,760 bytes, which is 5.11 GiB, so storage is not what stops you. The runtime is. Specifically, the documented path is a vLLM pooling server plus a Python package, and Ollama's registry has no clm tag today. In fact, our own probe of that registry finds nothing under contrastive-lm either. On top of that, nobody has published an end-to-end Mac test of either portable route.
The honest answer to "does it run on my Mac" has three parts. First, the weights fit comfortably on 16 GB. Similarly, the encoder half is portable to Apple Silicon through MLX or llama.cpp. Third, the last few percent of the pipeline is community work rather than an official path. In our experience fitting models to Apple Silicon, that third part is what decides whether a release is usable this month or next.
What CLM-8B actually is
To start with, the authors include researchers from Stanford and Nvidia, as VentureBeat reported at launch. The architecture is a frozen Qwen3-8B backbone with two small projection heads on top: a state head and an action head of about 20M parameters each. Specifically, the team trained both heads with a bidirectional InfoNCE loss. InfoNCE is a contrastive objective that pulls a positive pair together while pushing every negative apart. In practice, a state embedding is pulled toward the action that was taken and pushed away from every competing candidate.
There is no decoding loop at inference. Decoding is the token-by-token generation loop every chat model runs, and this architecture has none. Instead, the encoder embeds the state once, and the heads score each candidate action against that embedding. A softmax over those scores is then the answer distribution. A state embedding is simply the vector the frozen encoder produces for the state you hand it. In addition, the package exposes three typed questions plus a rank call:
Choice: pick among named options, with a probability per option.Noul: the probability that a statement is true.Score: a position on an ordered rubric.Engine.rank: score free-form candidates such as best-of-N answers, tool names or next moves.
For context, training ran in three stages, as the model card and the GitHub README describe them. Stage one used about 60M Nemotron question-answer pairs for pre-training. Stage two added about 30M synthetic hard negatives for mid-training. Subsequently, stage three used about 1M agentic trajectories for post-training. Notably, both encoder halves are cached separately, so action embeddings can be reused across states. Therefore, that caching is where the vendor latency claims come from.
The "System One" category is not new to this site. We covered the label and the open prior art in Jev's System One idea was open-sourced a year earlier, and Laya showed the same non-generative shape at 421M parameters. This article stays on the Mac question, and Ollama's decision lane covers the runtime side of the same trend.
The files, and what each one weighs
Encoder-only is the detail that trips people up. The community GGUF repo ships the Qwen3-8B encoder half quantised for llama.cpp. Quantisation is the trade of numeric precision for file size, and it is never free. Meanwhile, the heads are not quantised and not re-shipped there, because clm-serve fetches the reference head itself.
| File | Bytes | Size | What it is |
|---|---|---|---|
Qwen3-8B-BF16.gguf | 16,388,044,160 | about 16 GB | the unquantised reference encoder |
Qwen3-8B-Q8_0.gguf | 8,709,518,976 | 8.71 GB (8.11 GiB) | 8-bit K-quant, the build the quantiser recommends |
Qwen3-8B-Q6_K.gguf | 7,027,340,928 | 7.03 GB (6.54 GiB) | 6-bit, usable but flagged |
Qwen3-8B-Q5_K_M.gguf | 6,235,207,296 | 6.24 GB (5.81 GiB) | 5-bit, usable but flagged |
Qwen3-8B-Q4_K_M.gguf | 5,489,669,760 | 5.49 GB (5.11 GiB) | smallest 4-bit build, not recommended |
Qwen3-8B-Q4_K_M-outq2.gguf | 5,032,646,272 | 5.03 GB (4.69 GiB) | same quant, output projection at q2_K, pooling only |
MLX encoder model.safetensors | 4,730,780,891 | 4.73 GB (4.41 GiB) | unofficial MLX 4bit-g32 conversion, lm_head removed |
MLX heads CLM_v0.1-8B.safetensors | 75,552,331 | 75 MB | projection heads, distilled for that quantisation |
Reference head CLM_v0.1-8B.pt | 75,557,149 | 75 MB | the file clm-serve fetches from the parent repo |
For that, we pulled every byte count from the czl GGUF repo files, the mlx-community 4-bit conversion and the two Hugging Face file listings linked in Sources. Notably, the decimal labels in the table are the repos' own, with the binary unit beside them.
Does it fit your Mac?
The encoder is the floor, not the total. The heads add a 75 MB checkpoint, and the KV cache and the server add their own footprint on top. The KV cache is the per-request memory a server keeps for the token positions it is holding. Notably, a pooling server keeps far less of it than a chat server does. Even so, the pooled vectors are small but not free. Meanwhile, the one published Apple Silicon measurement is from the unofficial MLX port: 5.53 GB peak memory while answering, measured on an M3 Pro with 18 GB.
| Unified memory | What fits, in practice |
|---|---|
| 8 GB | Untested. The MLX port lists 8 GB Macs as untested, and the 4-bit GGUF encoder is 5.11 GiB before the KV cache |
| 16 GB | The MLX 4-bit build at a measured 5.53 GB peak, or a 4-bit GGUF. The 8.71 GB Q8_0 encoder is tight |
| 24 GB | Any 4-bit or 5-bit GGUF, and the Q8_0 encoder with room to spare |
| 32 GB | The same, plus the 6-bit and 8-bit files without juggling |
| 64 GB and up | The BF16 reference encoder, about 16 GB of weights, the closest thing to the exact configuration the card measured |
[ORIGINAL DATA] Read those tiers as our call, not a vendor statement. We derived them from the file sizes above, from the one published Apple Silicon measurement, and from the memory budget our own recommendation engine applies to every model in the catalog.
For a 16 GB Mac, then, the practical floor is the MLX route. It is the only Apple Silicon path with a published measurement, and its author recommends a 16 GB machine while calling 8 GB untested. Consequently, that route is where a reader on a small machine should start. The general rule for weight file versus runtime overhead is in how much VRAM for LLMs.
The published numbers, and who measured them
Every accuracy and latency figure in this release is the CLM team's own measurement. We did not reproduce any of them, and none of them was measured on a Mac.
- Zero-shot, across computer-use, gaming and tool calling: on par with Jev with up to 9x lower latency.
- With action caching and about 1k candidates: 13x faster than Jev.
- As a verifier, after fine-tuning the heads: 87.6 percent on Terminal-Bench 2.1 and 81.6 percent on DeepSWE, with verifier latency 4.1 to 5.7x lower than Jev.
- That evaluation covers 38 held-out DeepSWE tasks and 30 held-out Terminal-Bench 2.1 tasks, with candidate solutions generated by Opus 5 and Fable 5.
- The card states plainly that Jev falls below pass@1 as a verifier on those long-horizon tasks. In contrast, latency here was measured on an H100, not on Apple Silicon.
The GitHub README publishes the 87.6 and 81.6 figures, and both also appear on the model card. VentureBeat, which is where the Stanford and Nvidia attribution comes from, describes these as the team's zero-shot tests. Likewise, it records where Jev still won: all 30 WikiRacing tasks against CLM's 26, and the BFCL v4 tool-calling comparison. However, no protocol detail beyond that is published by either side, so treat the zero-shot table as vendor-reported.
Separately, two community builds published their own measurements, and each one flags a limit on its own conversion. That is worth reading before you download anything:
- The GGUF build checked every quant against a bf16 reference of the same encoder in the same runtime, over 23,926 scored questions. Q8_0 keeps the top option 0.9786 of the time and holds every decisive decision. Q4_K_M reaches 0.8429, loses 6.63 points of planner accuracy, and is flagged "not recommended" by its own author, past a 3.0-point budget. That is the one warning this article cannot skip: the smallest file in the repo is the one the person who built it tells you not to use.
- The MLX port compared itself to upstream's own server on 778 typed questions. It got top-option agreement of 91.4 percent, rising to 99.3 percent on questions where upstream was at least 70 percent confident, against upstream's 98.6 percent self-agreement. In addition, it calls itself an approximate build and says so on the card.
What it takes to run it on Apple Silicon today
In practice, three routes exist. Only two of them are portable to a Mac, and neither is the command in the model card.
1. MLX, the only Apple Silicon-specific route with published numbers. The port ships its own Python package inside the downloaded folder. Run pip install mlx mlx-lm transformers huggingface_hub, then hf download mlx-community/CLM-v0.1-8B-MLX-4bit --local-dir clm-mlx, then run the engine from inside that folder. In addition, its README notes the encoder weights are standard MLX Qwen3-8B, so general MLX tooling can load them. That said, the answer and rank behaviour is specific to this port and needs its own integration.
2. llama.cpp encoder plus clm-serve. The GGUF build documents a Metal-friendly server: llama-server -m Qwen3-8B-Q8_0.gguf --embedding --pooling last -c 16384 -np 8 -ngl 30 --port 8090, with clm-serve pointed at that endpoint. In particular, two details from that card matter. The context flag is the total across slots, so -c 2048 -np 8 leaves 256 tokens per slot and long states fail. Also, last-token pooling is baked into the files, which is why no runtime flag is needed: without it, the embeddings endpoint answers with a 400 error.
3. The documented vLLM command, vllm serve Qwen/Qwen3-8B --runner pooling --max-model-len 2048, is the path the authors measured on a GPU. Conversely, on a Mac the encoder has to come from MLX or llama.cpp instead, and that substitution is exactly the part nobody has published results for.
On the heads side, the question of a hidden download is worth settling, because it comes up. clm-serve fetches CLM_v0.1-8B.pt into ~/.cache/clm/ on first run, and that file is published in the parent repository: we found it listed there at 75,557,149 bytes, and clm-download prints the same cache path. Likewise, the heads are encoder-locked, meaning they only make sense against Qwen3-8B last-token-pooled embeddings. That contract is what both community builds reproduce.
What is missing is a published end-to-end answer test on a Mac. In contrast, the closest evidence in circulation is a cross-check in the GGUF card, where two independent implementations of the same encoder, llama.cpp and MLX, agree to a minimum cosine of 0.998680 over 4,448 texts. In addition, the two cases published on the model card reproduce in llama.cpp with the same argmax. In other words, the encoder side survives, yet that is not the same thing as confirming the whole pipeline answers identically on a laptop. We did not run it ourselves, so we are not going to present it as tested.
What a decision model is not
- It does not generate text. There is no chat template to feed a coding agent, and no
ollama run-style prompt to type into. The card's own words: CLM only scores the candidates you give it. - Its probabilities are relative to the candidate set you supply. If every option you pass is wrong, the model still has to rank them.
- The strongest verifier numbers need fine-tuned heads. The released checkpoint is the starting point, not the SOTA result.
- It is complementary to a reasoning model, not a replacement. The authors' own guidance is to generate and reason with large models, and to select, verify and monitor with CLM.
FAQ
Can I run CLM-8B with ollama run?No. The Ollama registry has no clm or contrastive-lm manifest, and its library does carry qwen3:8b as a generative tag, which is not the same artifact as a pooling encoder. The heads were trained on last-token-pooled embeddings, and the community GGUF build had to bake that pooling mode into the file for the embeddings endpoint to work at all. Therefore, CLM needs an embeddings server plus the head server, not a chat model.
Put plainly, less than you would guess. The unofficial MLX 4-bit build measured 5.53 GB peak while answering on an 18 GB M3 Pro, and its author recommends a 16 GB machine. Meanwhile, the 4-bit GGUF encoder is 5.11 GiB of weights before the KV cache, and the recommended 8-bit encoder is 8.71 GB, which is tight on 16 GB.
Which quant should I pick?In short, not the smallest one. In particular, the builder of the community GGUF flags Q4_K_M as not recommended, with top-1 agreement of 0.8429 and a 6.63-point planner accuracy loss, and points to Q8_0 as the closest to lossless. If memory is the binding constraint, the MLX 4-bit port is the alternative, at the cost of an approximate accuracy tier: 91.4 percent top-option agreement with upstream.
Where do the weights for clm-serve come from?From the parent model repository, into ~/.cache/clm/, as a 75,557,149 byte checkpoint that is listed in the repo. There is no private artifact behind it, and the community GGUF repo depends on clm-serve fetching it rather than re-shipping the heads.
Simply put, not end to end. The MLX port's throughput and memory figures are the only Apple Silicon measurements in circulation, and the GGUF repo's llama.cpp versus MLX cross-check is the closest thing to a Mac-side verification of the encoder. As a result, the pipeline as a whole, encoder plus heads plus answer parity, has no published Mac result.
Sources
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
The weekly local-AI refresh
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Have questions? Reach out on X/Twitter