By Peter · ModelFit · 2026-10-05

Kolibri 1 on Mac: 36GB Runs the 2-Bit Build, 64GB the 4-Bit, and Ollama Cannot Load It Yet

A Mac Studio on a dark desk beside a faintly etched Aleph Alpha plate, lit by cyan-teal light

Aleph Alpha released Kolibri 1 on 3 October 2026 under Apache-2.0: a German and English mixture-of-experts model with 78 billion total parameters that use 3.46 billion per token. There is no Mac path from the vendor. Aleph Alpha serves it with vLLM and its own plugin on data-center GPUs, and the published footprint is about 78 GB in FP8, with a stated minimum of one H200 or B200, or two A100 80GB cards. Everything that runs on a Mac therefore comes from community conversions published within a day of the release. The GGUF route needs a patched llama.cpp, and Ollama and LM Studio cannot load the architecture at all yet. Below is what each build costs, what the headline 1M-token context does to your memory, and how experimental these ports still are.

TL;DR: there is no vendor Mac path. 36GB of unified memory runs the 2-bit build, 48GB the 3-bit, 64GB the 4-bit. Ollama and LM Studio cannot load the architecture yet. Our own load attempt on a 16 GB Mac ends in swap, not in tokens.

What Kolibri 1 is

SpecValue
Released3 October 2026, Apache-2.0
Parameters78 billion total, 3.46 billion active per token
Experts384 per layer, 6 routed per token, plus 1 shared expert
Layers50, with a 4:1 sliding-window to full-attention ratio
Context1,048,576 tokens, with 262,144 or less recommended for serving
Weight formatFP8 (e4m3, 128x128 block scales)
ReasoningFour effort levels: none, low, medium, high
Tool callingOpenAI-style tools format
LanguagesGerman and English
Vendor hardware1x H200, B200 or B300, or 2x A100 80GB / H100 SXM5

The full spec list is on the model card. Mixture-of-experts is a design in which a router sends each token through a small subset of the model's experts instead of through all of them, and Kolibri 1 sends every token through 6 of the 384 experts in each layer. Two details of that design drive everything below. First, the model is sparse in compute and dense in memory: only a slice of the 78 billion weights is used per token, but all of them have to sit in memory, as the card states plainly. Second, the attention stack is mixed. Ten of the 50 layers run full attention, and the other 40 use a sliding window plus the current token. As a result, only those ten layers carry a KV cache that grows with the prompt.

Which Mac can run Kolibri 1?

Unified memory is the single pool Apple Silicon shares between the CPU and the GPU, and macOS caps how much of it a model may wire by default. That cap is why a 48GB Mac needs a manual tweak to hold a 35 GB build, and why the tiers below are about memory rather than chip speed.

Mac unified memoryBuild that fitsWhat it costs you
16 GBNothingThe smallest community build is 25.5 GB on disk, more than the whole machine.
24 GBNothingStill below the floor, and macOS keeps memory back for itself.
32 GBNothingClose, but below the 36 GB tier the guide sets for the 2-bit build.
36 GB (Mac Studio M5 Max)MLX 2-bit, 25.5 GBShort prompts only. The publisher quotes 26 GB with a short prompt and 30 GB at a 4k-token prompt.
48 GB (Mac mini M5 Pro, Mac Studio M5 Max)MLX 3-bit, about 35 GBRaise the GPU wired limit, and keep other large models closed. Publisher peak: 35 GB to 39 GB.
64 GB and up (Mac Studio M5 Max, MacBook Pro M4 Max)MLX 4-bit, 44.3 GB, or GGUF Q4_K_M at 47.5 GBThe closest thing to a comfortable setup. Publisher peak: 44 GB to 48 GB.
96 GB and up4-bit MLX with a long prompt, or the mixed 4/8-bit and 6-bit buildsRoom for context, not just weights.

What if you own a 24GB Mac mini or a 16GB MacBook Air? No build fits at any bit width, and that is the honest answer for most laptops. The tier mapping here follows from our analysis of the published file sizes against those limits, not from vendor guidance.

None of those rows has vendor support. The official minimum is one H200, B200 or B300, or two A100 80GB cards, because the FP8 weights alone take about 78 GB. In other words, every Mac route goes through a third-party quantization that Aleph Alpha did not publish, did not test and does not support.

What do the community builds actually weigh?

Three Apple Silicon Macs of increasing size lined up on a dark bench under cyan-teal light
BuildFile sizePublisher peak memory (short prompt / 4k prompt)Publisher perplexity, German / EnglishDownloads, 5 October 2026
MLX 2-bit, velaia25.5 GB26 GB / 30 GB14.01 / 17.46130
MLX 3-bit, velaiaabout 35 GB35 GB / 39 GB13.17 / 16.31153
MLX 4-bit, velaia44.3 GB44 GB / 48 GB12.77 / 16.4768
MLX 3-bit, here-be-dragons-ai33 GiB48GB Mac or moreNot published215
GGUF Q3_K_S, Eliasfpv2833.87 GBWeights are 31.54 GiB before buffersNot published675
GGUF Q4_K_M, Hob-forge47.5 GB64 GB of RAM, CPU testedNot published108

Sizes alone do not settle it, so which of those builds should you download? On a 64GB Mac, the 4-bit MLX build; on a 48GB Mac, the 3-bit. Some conversions also keep more precision where it matters, such as the mixed 3/6-bit build from here-be-dragons-ai.

The perplexity figures are the publisher's own measurement of its own MLX conversions, on the first 4,096 tokens of the German "Kolibris" and English "Hummingbird" Wikipedia articles. Treat them as roughly indicative, not as an independent benchmark: it is a small sample, from one person, about one set of conversions. Read them side by side anyway, because they show what the 2-bit build costs. Going from 4-bit to 2-bit moves German perplexity from 12.77 to 14.01 and English from 16.47 to 17.46 by the publisher's own count.

Meanwhile, that is the pattern we keep warning about in our note on 2-bit quants: below 3 bits, the size win stops being free. Therefore, on a 36GB Mac the 2-bit build is the only option, and on a 48GB Mac the 3-bit is already better on both sizes and scores. If you are buying hardware for this, the 4-bit build at 44.3 GB is the target, and everything below it is a compromise.

Even so, the download counts are the honest status signal. A few hundred people have pulled these files in two days: 675 for the Q3_K_S GGUF, 215 for the here-be-dragons 3-bit, 206 for the 6-bit MLX conversion by audreyt, 153 and 130 for velaia's 3-bit and 2-bit, 108 for Hob-forge's GGUF. That is early testing, not a supported path. Every one of these ports is a one-person effort and none of them is endorsed by Aleph Alpha.

Why can Ollama and LM Studio not load it yet?

The architecture identifier is kolibri1, and no released runtime ships support for it. Three concrete consequences:

  • Ollama and LM Studio. Both are built on llama.cpp, and both refuse the file until they include the new architecture. The Hob-forge GGUF card says it directly: Ollama, LM Studio and other llama.cpp-based apps will not load it yet. The run guide lists Ollama, LM Studio and Jan in the same column. The Ollama build on our own machine carries no kolibri1 architecture string in its binary.
  • llama.cpp. The GGUF conversions need a patch applied at upstream commit 836d571 before the build, then a llama-server or llama-cli built from it.
  • mlx-lm. Our installed mlx-lm 0.31.3 has no kolibri1 module, so a plain load of an MLX conversion stops with a "model type not supported" error. The velaia repos work around it with their own run.py launcher, which registers kolibri1.py from the model directory before handing over to the mlx-lm CLI. Support was proposed upstream in ml-explore/mlx-lm#1945 and was still open on 4 October 2026.

Practically: nothing here is a two-command install, and you should expect to build software, not just download a file. Our MLX versus Ollama comparison explains why the MLX route is the only one worth attempting on Apple Silicon today.

What does the 1M-token context cost in memory?

A glass-fronted GPU server rack in a dark colocation room with cyan-teal status lights and a coiled cable on the floor

KV cache is the per-token memory a model keeps for the keys and values it has already attended to. The card lists 1,048,576 tokens but recommends serving at 262,144 or less for efficiency and for complex tasks. On a Mac, however, the second number is the one to design around, because the weights are fixed and the KV cache is not. The MLX port in the velaia repos caches exactly what the config implies: 10 full-attention layers grow with the prompt, and the 40 sliding-window layers are held in a rotating cache capped at 513 tokens.

Do the arithmetic from the published config. Four key-value heads at a head size of 128 store 512 values per token per layer, and keeping a key and a value at bf16 makes that 2 KB per token per full-attention layer, so 20 KB per token across the 10 layers.

Context you actually fillKV cache at bf16
32,768 tokensabout 640 MiB
262,144 tokensabout 5 GiB
1,048,576 tokensabout 20 GiB

Those numbers fall out of our analysis of the published config, and they are arithmetic rather than a measurement. An FP8 cache halves them, which is what Aleph Alpha evaluates with on GPUs. The number to compare it against is the publisher's peak memory: the 2-bit build is quoted at 26 GB for a short prompt and 30 GB for a 4k-token prompt, which is already well above the 24 GiB weight file. Budget headroom, not the download size. Quality at length is a separate question. The guide reports 63.2 on RULER at the full million tokens against 85.4 on HELMET at 256K. That is the same pattern every long-context model shows: the far end of the window is not where you want to live.

Do the day-one quants actually answer?

Two data points, and neither is a benchmark.

The Eliasfpv28 Q3_K_S GGUF was started locally at a 4,096-token context, on an RTX 3060 12 GiB plus an Intel Arc Pro B60 24 GiB. Its card reports that the port answered a capital question and a multiplication correctly. The same card limits the claim: that is a functional check, not a quality evaluation, and reasoning mode, tool calling and longer contexts were not validated in that port.

By contrast, our own attempt ends in the same place. Downloading the smallest build, the 2-bit MLX conversion at 25.5 GB, completes without trouble. Loading it on the 16 GB M4 MacBook Pro in front of us does not. Before decoding starts, mlx-lm prints the mismatch: the model requires 24278 MB while the recommended maximum size for this machine is 12124 MB. The loader then spends several minutes materializing weights while macOS grows swap past 22 GB, and no first token ever arrives. We stopped it after four minutes. That is the 36 GB floor seen from the other side: it is not a preference, it is where the build stops thrashing.

Because of that, we have no decode speed of our own to report. The only speed figures in this article are the publisher's: 52 to 56 tokens per second on an M1 Max for all three MLX variants, which is consistent with a memory-bound mixture-of-experts model that activates 3.46 billion parameters per token.

If you want a working local model tonight instead, that is a different article. The honest advice for a 16GB or 24GB Mac is a 30B-class model, which is what our Mac Studio 64GB breakdown and the Qwen3.8-Flash-Next notes cover in detail.

Which Mac should you buy for Kolibri 1?

  • 64GB and up. The 4-bit MLX build at 44.3 GB is the target, with peak memory quoted at 44 GB to 48 GB. That is a Mac Studio M5 Max 64GB or a MacBook Pro M4 Max.
  • 48GB. The 3-bit MLX build at about 35 GB, after raising the GPU wired limit. Expect to close everything else while it runs.
  • 36GB. The 2-bit build at 25.5 GB, short prompts only, and a real quality cost against the 4-bit.
  • 16GB to 32GB. Not this model at any quantization. The weights do not fit, before any context.

Two practical warnings. First, these are unvalidated ports of a two-day-old architecture, and one of them is a 675-download GGUF from an individual account. Second, the sane move is to wait: once kolibri1 lands in mlx-lm upstream and in a llama.cpp release, the same weights load with a normal command and Ollama catches up. Until then, treat Kolibri 1 on a Mac as an experiment, and check the general memory rules before you size a purchase around it.

FAQ

Can a 16GB Mac run Kolibri 1?

No. The smallest community build is 25.5 GB on disk at 2 bits, and macOS needs part of that 16 GB for itself. Loading it on our 16 GB M4 MacBook Pro ends in swap, not in tokens. The practical floor is a 36GB Mac with the 2-bit build.

Does Ollama support Kolibri 1?

Not yet. Ollama is built on llama.cpp, and the kolibri1 architecture is not in either of them. The GGUF cards say the file will not load until support is added, the current Ollama build carries no kolibri1 string, and the alternative is a patched llama.cpp build from a specific commit.

Is the 2-bit MLX build good enough?

Only if 36GB is all you have. The publisher's own numbers put German perplexity at 14.01 and English at 17.46 at 2 bits, against 12.77 and 16.47 at 4 bits, and our standing advice is to stay at 3 bits or above when the memory allows it. On a 48GB Mac, take the 3-bit build.

Can I use the full 1M context on a Mac?

Not realistically. Aleph Alpha recommends serving at 262,144 tokens or less, and the KV cache for the ten full-attention layers costs about 640 MiB at 32,768 tokens, about 5 GiB at 262,144 and about 20 GiB at the full million, on top of the weights. On a 36GB or 48GB machine the weights already eat most of the budget.

How good is Kolibri 1 compared with Qwen?

Aleph Alpha's own comparisons put it ahead on math, with 87.5 on the German AIME 2025 against 82.9 for Qwen3.6-35B-A3B. They put it behind on tool calling, 61.4 against 67.2 on BFCL v4, and on long-context retrieval, 64.5 against 70.8 on LongBench Pro. Those runs are the vendor's, not an independent evaluation. The full breakdown notes that where Qwen wins, for example on tool calling, it wins with less than half the stored weights.

Sources

What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

Have questions? Reach out on X/Twitter