Hayder Tirmazi published a set of optimisations to prompt lookup drafting in llama.cpp on 26 September 2026. He measured them on an Apple M4 Pro with 14 cores and 48 GB of memory, and he reports up to 42x faster drafting and up to 2.6x less memory. The measurements are real, and the code does what the article says. However, none of it is in the llama.cpp you already installed, and it has almost nothing to do with tokens per second. So we rebuilt the fork on an Apple Silicon Mac, confirmed the direction of every claim, and ran the test nobody published: real generation, base versus fork, same model and same prompt.
So what does a 42x drafting number actually buy you on a Mac? The short version: this buys faster startup and much lower RAM for the static n-gram cache. It does not buy generation speed, and it lives in a fork whose pull requests are all still open.
What the 42x measures, and what it does not
Prompt lookup decoding is a speculative method that guesses the next few tokens from n-gram statistics instead of running a second model; other authors call it n-gram speculation. The draft loop is the part that does the guessing. Moreover, llama.cpp keeps three n-gram caches: a context cache built from the tokens currently in context, a dynamic cache carried across runs, and a static cache built offline from a text corpus with the llama-lookup-create tool. An n-gram cache is a lookup table that maps short token sequences to the tokens that usually follow them, so a hit is a cheap guess and a miss costs nothing. The loop scores candidate tokens for n-grams of size 1 to 4, checks them against hard-coded thresholds, and hands the survivors to the target model for verification. Prompts that quote or copy their own context are where this pays off.
The headline number is the cost of that draft loop. Tirmazi measures it with llama-lookup-stats, a tool that reads a text file, treats its tokens as if they were a model output, and replays the drafting loop over them. That replay is a CPU-only experiment: no model inference happens inside it, and a GGUF file is only needed for the tokenizer. As a result, the article never mentions Metal, GGUF model choice, or throughput. Its metric is microseconds per drafted token inside one CPU data structure, not tokens per second.
His test rig, as stated in the article, is an Apple M4 Pro with 14 cores and 48 GB of memory, a 4096-token context, and WikiText-103 static caches at 25, 50, 100 and 200 MB plus the full corpus at about 541 MB, with a median of 3 runs. A corpus size of 0 means no static cache. That is the common case for most readers, so if you run chat prompts rather than document copies, the 541 MB row is not your row.
The gains, step by step, with the configuration attached
Every figure below is the author's own, from his own branch, and we have not seen an independent reproduction. Note that his numbers mix two different setups: with a static cache and without one. Therefore, keep them separate when you read the ranges.
| Change | Effect, no static cache | Effect, with a static cache |
|---|---|---|
| Read inner maps by reference instead of copying | drafting 4.5x to 25.6x faster by corpus size | same change, same range |
| Outer map to ankerl unordered_dense segmented map | drafting 1.02x to 1.13x faster | cache load 1.41x to 1.65x faster, cache RAM 1.07x to 1.11x lower |
| Inner map to a sorted vector | drafting 2.09x faster | drafting 1.19x to 1.25x faster, peak memory up to 1.97x lower |
| Static cache backed by a constmap | not applicable | loads 6.32x to 16.12x faster (3.76 s to 0.23 s at 541 MB), cache RAM 463 MB for a 467 MB file, peak memory 1.71 GB to 1.31 GB, drafting 1.06x to 1.20x faster |
| Daniel Lemire early threshold check (fork PR 12, still open) | drafting up to 1.9x faster | drafting up to 4.2x faster |
Two details matter for reading that table. First, the 1.71 GB to 1.31 GB peak memory figure comes from a comparison against the previous step on the same branch, not against upstream master. Consequently, any summary that quotes a fall from 3.36 GB to 1.31 GB is quoting a figure that does not appear in the article text. Second, the constmap step is the only one that touches the static cache file itself, which is why it produces both the large load-time gain and the large RAM gain. A constmap is a flat, sorted key-value structure that stores each key once and answers lookups without allocating on the read path, which is exactly the shape the static cache needs. The 2.6x less memory claim in the TL;DR, and the 140x total after Lemire's patch, are composites of the rows above.
How do you know which rows apply to your workload? Count how much of your output copies text that is already in the prompt. For example, a document Q&A bot that feeds a long contract into context hits the n-gram tables far more often than a short chat turn does, so a reader in that situation lands on the static-cache rows. A reader chatting about the weather lands on the row with no cache at all.
The threshold change in Lemire's patch is small. It checks whether the most frequent follower of an n-gram can pass the acceptance threshold before scoring all the candidates, and it takes the inner map entries by reference instead of by value. In total, the patch adds 15 lines and removes 7 lines in common/ngram-cache.cpp.
Is any of it in the llama.cpp you installed?
No. The work lives in jadidbourbaki/llama.cpp, a fork created on 24 September 2026 whose master tracks upstream at commit 84e76d8a. Every pull request implementing the optimisation is open and unmerged: the copy fix (#2), the unordered_dense outer map (#5), the constmap static cache (#7), the sorted-vector inner map (#10), and Lemire's threshold check (#12). Three earlier attempts (#1, #3, #6) are closed and unmerged. We re-queried the GitHub API on 28 September 2026 and found no pull request for this work on ggml-org/llama.cpp. The only PR matching a prompt-lookup search upstream in the same window is #28992, a server prompt cache fix from 16 September, unrelated to drafting.
There is a reason nothing was filed upstream, and it is not technical. Asked about the work on Hacker News, the author said plainly that he cannot open a pull request or an issue in the llama.cpp repository because of an interpersonal dispute, and he pointed readers to a LocalLLaMA thread. That thread carries pushback as well. Whatever the merits, the practical consequence is the one that matters here: this code has no path into your installation unless a third party carries it upstream.
The Hacker News thread on the post sits at 71 points and 11 comments. The single technical question in it asks how much this speeds up inference in tokens per second. It is unanswered, and that question is the whole article below.
What we measured on Apple Silicon
Does the direction of every claim hold on a machine that is not the author's M4 Pro? That is the question we set out to answer. So we cloned both trees and built them here: upstream master at 84e76d8a as the baseline, and the fork with Lemire's PR 12 applied as the optimised build. Both configure and compile cleanly with CMake on Apple Silicon, Metal enabled, with no source edits and no warnings that stop the build. The changed files sit in common/, vendor/ and examples/lookup/, and no backend or Metal kernel code changed. Therefore, the optimisation is architecture-independent by construction. In practice, we measured it on an M4 with 10 cores and 16 GB, not the author's M4 Pro, and with a 269 MB WikiText-103 slice, roughly half of his largest corpus.
The replay benchmark below reports medians of 3 runs, in microseconds per drafted token:
| Static cache corpus | Upstream master | Fork plus PR 12 | Drafting speedup | Cache load, master to fork | Acceptance rate |
|---|---|---|---|---|---|
| none | 15.94 us | 0.80 us | 19.9x | no load step | 30.4% to 30.8% |
| 25 MB | 68.33 us | 1.14 us | 59.9x | 384 ms to 32 ms (11.9x) | 32.08% to 32.15% |
| 100 MB | 124.08 us | 1.59 us | 78.0x | 1231 ms to 91 ms (13.5x) | 34.11% to 34.18% |
| 269 MB | 190.67 us | 1.61 us | 118.4x | 2658 ms to 159 ms (16.7x) | 35.31% to 35.36% |
The direction of every claim reproduces, and our speedups are larger than the headline 42x because drafting latency grows with cache size on the baseline and stays flat on the fork. Importantly, acceptance rates match within 0.07 percentage points, which is the author's own point about not changing what the decoder does.
The memory claim reproduces cleanly, and it is the most useful part of the change. We measured the whole process, including a 2B model loaded for tokenisation. On upstream master the static cache costs a little more than 5x its file size in RAM, and on the fork it costs about 1x:
| Static cache corpus | Cache file | Peak RAM on master | Peak RAM on fork | Saved | RAM per byte of cache file |
|---|---|---|---|---|---|
| none | 0 | 1.67 GB | 1.62 GB | 0.05 GB | n/a |
| 25 MB | 49.6 MB master, 47.2 MB fork | 1.92 GB | 1.67 GB | 0.26 GB | 5.13x master, 1.00x fork |
| 100 MB | 143.6 MB master, 137.5 MB fork | 2.42 GB | 1.76 GB | 0.66 GB | 5.22x master, 1.00x fork |
| 269 MB | 299.1 MB master, 287.3 MB fork | 3.26 GB | 1.91 GB | 1.35 GB | 5.32x master, 1.00x fork |
Building the cache in the first place is the other cost, and the fork does not change it much: 70.3 s on master and 59.7 s on the fork for a 269 MB corpus, at a peak of about 7.6 GB of RAM in both cases. That is the step that actually needs memory on an 8 GB or 16 GB Mac. So on a 16 GB machine, the cache build is what you plan around, not the cache load.
Does 42x drafting make generation faster? Measured, no
The article publishes no tokens-per-second figure, and the tool it benchmarks cannot produce one. For example, llama-lookup-stats replays a text file and never runs a model. So we built llama-lookup in both trees and ran real generation: same Qwen3.5-2B Q4_K_M model, same 5200-character passage with a verbatim copy instruction, temperature 0, fixed seed, 512 tokens, 3 runs each, same static cache size, 4096-token context.
| Configuration | Master | Fork plus PR 12 | Difference |
|---|---|---|---|
| 269 MB static cache | 87.4 t/s | 82.5 t/s | -5.7% |
| 25 MB static cache | 85.8 t/s | 80.2 t/s | -6.6% |
| no static cache | 81.5 t/s | 81.2 t/s | -0.4% |
Generation speed is unchanged within the spread, and if anything slightly lower for the fork. One honest caveat: acceptance rates diverged between the two builds in these runs (87.8% against 94.4% at 269 MB, 84.5% against 73.0% at 25 MB, 77.6% against 70.3% with no cache). That divergence contradicts the author's statement that acceptance is identical, and it makes the end-to-end comparison less well matched than we would like. However, the direction of the conclusion does not depend on it.
The reason is arithmetic, not opinion. So why does a 42x win on the draft loop vanish in real generation? In the heaviest replay configuration we ran, upstream master spends 190.67 us drafting per drafted token. At 35.3% acceptance, that is about 0.54 ms of drafting per accepted token, against 11.4 ms for one generated token on this machine, roughly 4.7% of the time. The fork brings that to 0.04%. In the real generation run with a copy-heavy prompt, drafting took 0.21% of per-token time before and 0.02% after. There was never enough time there to win back.
What you do feel is startup. Loading the 269 MB static cache drops from 2658 ms to 159 ms, and the 25 MB cache drops from 384 ms to 32 ms. On a server that reloads a cache per process, or a script you restart constantly, that is a visible change. By contrast, a single long chat session will not notice it, because drafting was never the bottleneck there.
Compatibility by unified memory tier
The drafting code is CPU-side common code with no backend dependency, so it runs on any Apple Silicon Mac. What changes by tier is how much static cache you can afford. Which row is yours? The table below uses our measured memory multipliers and a 2B to 4B model at 4096 tokens of context. For example, a 16 GB M4 in the second row can hold a 269 MB corpus cache with a 2B model, and anything much larger starts to squeeze the model itself.
| Unified memory | Upstream master | Fork plus PR 12 |
|---|---|---|
| 8 GB | Keep the static cache at a 50 MB corpus or below, or skip it. A 269 MB corpus cache costs about 1.6 GB of extra RAM here and pushes a 4B model into swap. | A 269 MB corpus cache costs about 0.29 GB. Up to roughly a 1 GB corpus stays comfortable, an estimate from the measured 1.0x multiplier, not a measurement. |
| 16 GB | A 269 MB corpus cache fits with a small model but leaves little room. Measured peak 3.26 GB with a 2B model loaded. | Same workload peaks at 1.91 GB, leaving about 1.35 GB more headroom. |
| 24 GB | Comfortable up to a few hundred MB of corpus. | Comfortable well beyond it, limited by cache build time rather than RAM. |
| 36 GB and above | The static cache is not the constraint at any corpus size tested here, anyone can run this. | Same, with the caveat that the cache build step peaks at about 7.6 GB regardless of build, which is harmless on this much memory. |
One warning that no table covers: the static cache file format changed. A cache built by upstream master makes the fork abort with a magic number assertion, and a cache built by the fork makes master abort on a count assertion. We triggered both failures here. So rebuild the cache when you switch builds, and remember that the file is not portable to a stock release build.
How to evaluate this without betting your setup
How do you test an unmerged patch without risking a working install? Build it in a scratch directory, not over an installation you depend on:
git clone --branch ngram-cache-constmap https://github.com/jadidbourbaki/llama.cpp
cd llama.cpp && git fetch origin pull/12/head:pr12 && git checkout pr12
cmake -B build -DLLAMA_CURL=OFF -DLLAMA_BUILD_TESTS=OFF && cmake --build build -j
Then build a static cache from a corpus of your own with llama-lookup-create, and compare llama-lookup-stats medians before and after. If your workload is a chat with short prompts, the honest expectation is that nothing changes except startup, because prompt lookup only drafts well when the output copies text that is already in context. By contrast, if you run long contexts with quoted documents, that is where the Metal side of the stack still dominates your throughput, and this optimisation does not touch it.
If what you want is generation speed rather than drafting elegance, the releases that shipped Metal and kernel work are the ones that move your tokens per second: see our notes on llama.cpp 0.5.0 MoE fusion and on Metal kernels for GGUF models.
FAQ
Is this optimisation in any llama.cpp release?
No. As of 28 September 2026 all five pull requests are open in the fork jadidbourbaki/llama.cpp, none are merged, and GitHub search shows no equivalent pull request on ggml-org/llama.cpp. The thresholds the article describes match release b11182 of upstream, so the code path it optimises is the one you are running. However, the optimisation itself is not in that release.
Did the author publish a tokens per second figure?
No. The benchmark tool replays a text file as if it were model output, so it never runs a model and cannot produce a throughput number. The only tokens-per-second question in the 11-comment Hacker News thread is unanswered. Meanwhile, our own end-to-end runs on an M4 show generation speed unchanged.
How much RAM does the static n-gram cache really use?
On upstream master, a little over 5x the cache file size: 1.59 GB of extra RAM for a 299 MB cache at a 269 MB corpus, measured. On the fork with the constmap static cache, about 1.0x, so 288 MB for the same corpus. If RAM is your reason to read this post, that is the figure to take away.
Does it work on M1, M2 and M3 Macs too?
The changes are confined to common/ngram-cache.cpp, common/ngram-cache.h and vendored headers, with no Metal, BLAS or other backend code touched, so nothing about it is M4-specific. We built and measured it on an M4. We have not tested M1, M2, M3 or an Intel Mac, and the author measured on an M4 Pro only.
Should I switch my daily driver to this fork?
Not for generation speed. Switch only if you specifically want faster static cache loading and lower cache RAM. If you do switch, rebuild the cache file first, and go in knowing that the format is incompatible with stock builds and that the code is unmerged, unsupported, and blocked from upstream by a dispute rather than by review.
Sources
- Hayder Tirmazi, "42x faster prompt lookup drafting in llama.cpp", 26 September 2026: https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/
- Fork: https://github.com/jadidbourbaki/llama.cpp (created 24 September 2026, master at 84e76d8a)
- Daniel Lemire threshold check, PR 12, open and unmerged: https://github.com/jadidbourbaki/llama.cpp/pulls/12
- constmap reference implementation used by the static cache step: https://github.com/lemire/fastconstmap
- Hacker News discussion, 71 points, 11 comments: https://news.ycombinator.com/item?id=49859982
- Own measurements, 28 September 2026: Apple M4, 10 cores, 16 GB, macOS, upstream master 84e76d8a versus fork PR 12 head f764f323, Qwen3.5-2B Q4_K_M for tokenisation, WikiText-103 test split for replay, medians of 3 runs.
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
The weekly local-AI refresh
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Have questions? Reach out on X/Twitter