Liquid AI published the open d1 decision models on 2026-10-07, and unlike most multi-billion-parameter releases this one fits the laptop you already own. d1-3B packs 3.12 billion parameters, answers typed questions with calibrated probabilities instead of writing tokens, and had a GGUF build on Hugging Face the day before the announcement. The useful question is not whether a 3B decision model is interesting. It is whether it runs on your Mac, and what it replaces in a local pipeline.
Does d1-3B run on your Mac?
Yes, and the footprint is small. A decision model is a model that reads a state and a set of named questions once, then returns a probability for every allowed answer without generating text. That design removes the token-by-token decode, so what you size is just the weights plus the vision projector when you send images.
The GGUF repo lists three weight files:
- Q4_K_M at 1,674,456,672 bytes (about 1.7 GB)
- Q8_0 at 2,874,781,280 bytes (about 2.9 GB)
- a 16-bit build at 5,403,160,160 bytes (about 5.4 GB)
The vision projector is a separate file, 583,109,728 bytes for the Q8_0 version (about 0.85 GB for the F16 one). The tiers below are our call, derived from those file sizes plus the RAM macOS keeps for itself, not a vendor statement. For the general per-parameter rule behind this math, see how much VRAM local LLMs need.
| Unified memory | Q4_K_M (about 1.7 GB) | Q8_0 (about 2.9 GB) | 16-bit (about 5.4 GB) |
|---|---|---|---|
| 8 GB | Fits comfortably | Workable | Does not fit well with a long state |
| 16 GB | Fits comfortably | Fits comfortably | Workable |
| 18-24 GB | Fits comfortably | Fits comfortably | Fits comfortably |
| 32 GB and up | Fits comfortably | Fits comfortably | Fits comfortably, room to batch |
Get told when a better model fits your MacBook Air M5 16 GB
One email when a new open-weight model is a better fit for your machine, plus the Thursday weekly.
For: MacBook Air M5 16 GB
Unified memory is the single pool of RAM an Apple Silicon Mac shares between its CPU and its GPU, so the weights and the projector both come out of it. Add about 0.85 GB when you plan to send images. There is no generation phase, so nothing grows while you wait, and the model does not stream tokens back. The whole lane sits well under the memory floor that most 30B-class chat models need.
What did Liquid actually ship, and when?
The family landed across three days. The d1-3B repository and d1-omni-600M were both created on 2026-10-05, the d1-3B-GGUF repository on 2026-10-06, and a third head, d1-3B-w8a8, on 2026-10-07. Liquid announced the release on its blog on 2026-10-07. So the weights, the quantized build, and the announcement did not arrive together, and the GGUF predates the blog post by a day.
Two models are open-weight today:
- d1-3B, a 3.12B multimodal decision model built on LFM2.5-VL-3B, with a SigLIP2 NaFlex 400M vision encoder, a 32,768-token context, and a 128,000-token vocabulary.
- d1-omni-600M, an experimental model that takes text with an image, or text with an audio clip of up to 30 seconds. Liquid states it does not report inference numbers for d1-omni-600M because it is an early research release under active development.
Both are decision models from the first line of the card, and both say plainly that they are not chat models and do not write text.
What is d1, and what is it not?
A conventional LLM writes an answer token by token and your code parses the prose. d1 reads the state and the questions in one forward pass and returns the distribution over the allowed answers. Every response reports output_tokens: 0. That single design choice is the whole speed argument, and it is also the limit of what the model does.
Three question types are defined on the card:
noul: a yes or no question, answered as P(yes), a probability between 0 and 1.choice: one label from named options, returned with a full probability distribution and a confidence value.score: a probability-weighted position on an ordered rubric of 2 to 10 levels.
Several questions can share one state in a single call, which the card highlights as the token-saving case. The recommended uses are routing and triage, moderation, intent and topic classification, extraction checks, reranking, judge scoring, agent guardrails, and visual inspection. What it is not is a general assistant: it does not hold a conversation, does not summarize, and does not write a sentence you can hand to a user without a template. If you were hoping to swap a chat model out, this is the wrong family.
How do you run d1-3B on a Mac?
There are two paths, and they are not the same.
The first is the reference loader. The card ships its own code and asks for transformers>=5.14, and the sample selects the Apple GPU through mps when no CUDA device is present. That is the path behind Liquid's own Mac timing.
The second is llama.cpp, which is what most people will want. The GGUF card gives the exact command:
llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0
You then post to the /v1/systemone endpoint on port 8080. That endpoint is documented in the llama.cpp server README, and it shipped in pull request #29818, merged 2026-10-02 and first released in build b11361. Support for the d1-3B architecture itself is a later change: pull request #30110 added the model on 2026-10-07, after the newest tagged build at the time of writing, b11476 (2026-10-07T14:55Z). Read that carefully before you pull a binary: the endpoint is available, but a tagged build that loads d1-3B specifically is not in the release list yet, so plan on a source build from master on or after 2026-10-07 or a newer nightly.
On Ollama, the honest answer is no. Ollama has its own decision lane behind /v1/systemone with the nimble and tev1 models, which we covered in the Ollama decision-model release, and d1-3B is not published there. Do not assume ollama run works on this family. The GGUFs are llama.cpp files.
What do the benchmark numbers actually say?
Read the attribution before the digits. Every figure below is a vendor measurement. Liquid states it scored d1 with the official scorer rather than submitting to the leaderboard, and the competitor rows come from the public leaderboard v0.2.1. There is no independent reproduction of any d1 number, and we did not run the model ourselves, because ModelFit runs no benchmarks.
| Model | Size | Decision Index 0.2.1 |
|---|---|---|
| Winnow-12B | 12B | 50.02 |
| d1-3B | 3B | 48.57 |
| Decider 35B-A3B | 36B | 47.11 |
| JPT-9B | 9.7B | 46.89 |
| Decider 4B | 4.7B | 40.70 |
| d1-omni-600M | 587M | 15.95 |
The headline claim is that 48.57 makes d1-3B the best decision model under 10B on the Decision Index 0.2.1, ahead of every 4B and 9B model and of Decider 35B-A3B at 47.11. It beats the 36B Decider by 1.46 points on the total while trailing it on Knowledge. On the card's per-category split, d1-3B leads on Tools (74.5) and Arts (36.3) and is weakest on Knowledge (23.8). The small d1-omni-600M sits at 15.95 on the same index, which is the honest signal that the 600M class is not a stand-in for the 3B.
On vision, the card reports 74.1 on 11 public image benchmarks against 73.9 for its LFM2.5-VL-3B base, so the post-training did not cost image quality. Treat that figure as a vendor claim with no third-party repeat.
The latency table is where the Mac number lives. Liquid measures Apple Silicon on an M5 Pro through mps, not through llama.cpp:
| Path | One question | Three questions, one pass | 3.4k-token state | 384 px image |
|---|---|---|---|---|
| Apple M5 Pro (mps) | 30 ms | 41 ms | 640 ms | 62 ms |
| NVIDIA RTX 4090 (bf16) | 8 ms | 21 ms | 102 ms | 17 ms |
| AMD MI325X (bf16) | 9 ms | 14 ms | 44 ms | 18 ms |
The 30 ms is Liquid's own figure, measured on an M5 Pro under the reference loader, warm, one request at a time. It is not an independent measurement, it is not a llama.cpp measurement, and it is not a number we produced. The realistic reading is that d1-3B answers a single question in tens of milliseconds on a recent M-series chip once the shapes are warmed up, and that a long state dominates: 640 ms for the 3.4k-token example is a different order of magnitude from the 30 ms warm call. The card also notes that the first call with a new shape pays for kernel compilation, so warm the shapes you serve.
Can you ship d1 in a commercial product?
Not under Apache 2.0. The weights use the LFM Open License v1.0, and the license text defines a Threshold of annual revenue of 10 million United States dollars ($10,000,000) or more. Section 5 conditions commercial-use rights on you not exceeding that threshold, and states that any commercial use by a legal entity above it is not licensed under the agreement. Qualified non-profits are exempt for non-commercial or research purposes. Below the threshold, the grant is perpetual, worldwide, and royalty-free, with the usual attribution and notice conditions. So the practical line is: free to use and ship commercially under 10 million dollars in annual revenue, and a separate agreement above it. That is a real constraint for a funded startup, and it is not the same permission set as the Apache-2.0 decision models we compared in the Kev family on Apple Silicon.
What does d1 replace in a local pipeline?
A decision model replaces the classification call, not the assistant. The concrete cases from the card map onto common local jobs:
- Routing a support ticket to billing, technical, or fraud before any chat model sees it.
- Triaging inbound text by intent or urgency, scoring each request on an ordered rubric.
- Reranking retrieved passages for a local RAG pipeline.
- Acting as a guardrail that returns a probability instead of a paragraph.
- Inspecting an image for a yes or no before you spend a vision-language model on it.
Compared with a chat model, you trade prose for a letter and a probability. You also get a number you can threshold, so uncertain cases route to a human. That is the same bargain the earlier Laya decision encoder and the Ollama decision lane strike, and it is why 3B is enough here: the task is a single distribution, not a chain of tokens. If your pipeline already runs an LFM model through Ollama, the base model is the same family that LFM2.5 on Apple Silicon covers, though d1 is post-trained for decisions rather than chat.
FAQ
Is d1-3B a chat model I can replace my local LLM with?No. The card states it is not a chat model and does not write text. It returns typed answers (yes or no, a chosen label, a score) with probabilities, in one forward pass with zero output tokens. Use it for routing, triage, classification, reranking, and guardrails, not for conversation.
What Mac do I need for d1-3B?Any Apple Silicon Mac with 8 GB of unified memory can hold the Q4_K_M build at about 1.7 GB, and the Q8_0 build at about 2.9 GB is workable there. A 16 GB Mac runs any quant comfortably, and you add about 0.85 GB if you send images. Those tiers are our call from the file sizes, not a vendor figure.
Does Ollama run d1-3B?No. Ollama serves its own decision models, nimble and tev1, behind the same /v1/systemone endpoint. d1-3B is distributed as GGUF and documented for llama.cpp. Do not expect ollama run to fetch it.
It is Liquid's own measurement, reported on the d1-3B card for an Apple M5 Pro under the reference mps loader. No independent test has reproduced it, and we did not run the model. The card also reports 640 ms for a 3.4k-token state, so the per-question time depends heavily on how much state you send.
Below 10 million dollars in annual revenue, the LFM Open License v1.0 grants a perpetual, royalty-free commercial use. At or above that threshold, commercial use is not licensed under the agreement and needs a separate deal with Liquid AI. The weights are not Apache 2.0.
Sources
- Liquid AI open d1 model card, d1-3B: https://huggingface.co/LiquidAI/d1-3B
- Liquid AI d1-3B GGUF card and llama.cpp run path: https://huggingface.co/LiquidAI/d1-3B-GGUF
- Liquid AI d1-omni-600M model card: https://huggingface.co/LiquidAI/d1-omni-600M
- llama.cpp server README,
/v1/systemoneendpoint: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md - llama.cpp pull request #29818, add
/v1/systemone: https://github.com/ggml-org/llama.cpp/pull/29818 - llama.cpp pull request #30110, add the d1-3B model: https://github.com/ggml-org/llama.cpp/pull/30110
- LFM Open License v1.0 text: https://huggingface.co/LiquidAI/d1-3B/raw/main/LICENSE
Don't miss the next one for your MacBook Air M5 16 GB
This article covers one moment. Get one email when a newer open-weight model fits your machine better, plus the Thursday weekly with what changed.
For: MacBook Air M5 16 GB
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
Have questions? Reach out on X/Twitter