By Peter · ModelFit · 2026-10-01

Kev Gives Your Mac a Jev-Style Decision Model: Kev-0.8B Fits Any Apple Silicon, Kev-27B Wants 96 GB

A silver MacBook Pro on a dark desk lit by cyan-teal rim light, with a small metal token placed beside it.

Kev is a four-model family of Jev-style decision models published by Jared Palmer under Apache-2.0, and the Mac question has a short answer and a long one. The short answer: Kev-0.8B is documented for any Apple Silicon Mac, Kev-4B and Kev-9B are documented for a 32 GB Mac, and Kev-27B is documented for a 96-128 GB Mac that the author has not measured. The long answer is that the project's two sizing tables disagree with each other, the only hardware figure attached to Kev-27B on Apple Silicon is an expectation rather than a reading, and the first hard memory numbers for Kev-4B on a Mac come from a third party's unofficial quantised build. Everything below is either the author's number or that third party's number, and every page it comes from is linked.

Key takeaways: Kev-0.8B is the only size the project claims for any Apple Silicon Mac. Kev-4B is the recommended default, yet only an unofficial 8-bit MLX build has a published Mac memory reading. Kev-27B's 96 GB figure comes from its weight size, and nobody has published a Mac run at 27B.

What Kev actually is

A decision model is a network that reads one state plus a set of typed questions and returns labels and probability distributions, with no tokens generated. Kev serves TypeSafe's public /v1/systemone contract. Specifically, choice questions pick among named options, noul questions return a yes/no probability, and score questions place the state on an ordered rubric. All questions in a request share the same state but cannot read each other. Similarly, the probabilities come back calibrated, because each checkpoint ships with a temperature fitted on held-out data. Calibration is the fit between a stated probability and how often the model is right at that probability. The API shape is deliberately the one TypeSafe's hosted Jev uses, so the TypeSafe Python SDK works against a Kev server unchanged.

The family is two architectures wearing one name. Kev-0.8B, Kev-4B and Kev-9B are rank-16 LoRA adapters plus a small pointer head on frozen Qwen3.5 base models, so the released tarballs carry an adapter rather than full weights. A pointer head is the small module that scores each option's closing token against the question's final token, and a softmax over those scores becomes the answer distribution. Kev-27B starts from Qwen's post-trained Qwen3.8-27B release and fine-tunes every weight, so it ships as a full bf16 checkpoint instead. None of the Kev outputs were trained on Jev outputs. In fact, the whole family landed in one week: the first checkpoint (Kev-0.5B on Qwen2.5-0.5B) on 2026-09-17, and then the current four sizes with the kev-family release on 2026-09-20.

This is the same lane we already covered from the other direction. The label itself and the open prior art are in Jev's System One idea was open-sourced a year earlier, while Laya showed the non-generative shape at 421M parameters on Apple Silicon. Meanwhile, the runtime side of the trend is in Ollama's decision lane.

Does the MLX path work on a Mac today?

Yes, as a documented path with published latencies, and with one caveat that matters. The documented local command is a two-step: uv sync --extra serve, then python -m kev.serve --run jaredpalmer/kev-4b. The README states that this picks CUDA or ROCm when a GPU is present and MLX on Apple Silicon. In addition, the first run downloads the adapter and the base model from the Hub. MLX is Apple's array framework, and the serving code uses it through mlx-lm, so there is no CUDA fallback to arrange by hand.

The author's Mac table measures five questions about a roughly 270-token text on an M5 with 32 GB:

ModelNew textSame text again
Kev-0.8B149 ms28 ms
Kev-4B721 ms136 ms

Those come from the same README that serves the example response labelled bf16 on an Apple M5, and the server runs in bf16 on Macs as well as GPUs. Two model cards say the same thing in older words. The Kev-4B card still describes the PyTorch MPS fallback, where a five-question request takes 0.78 s in bf16 on an M5. The Kev-9B card says kev.serve runs that checkpoint through MLX on Apple Silicon, because the DeltaNet kernels have no MPS implementation. That 0.72 s to 0.78 s range for the same workload lines up. However, the 4B card's advice to use the older @qwen3 revision for low Apple Silicon latency predates the MLX backend. In other words, it no longer matches the README.

The honest framing is that the MLX path works and is the default. That said, the numbers published for it are latency figures rather than memory readings. We checked both published tables for a resident-memory figure on a Mac and found none.

The compatibility table, with the disagreement made explicit

Close-up of an open Mac laptop's aluminium edge with cyan-teal light seeping from the display gap on a dark background.

Two tables in the project disagree about the floor, and the difference is not cosmetic. The README's model table names a Mac: any Apple Silicon Mac for Kev-0.8B, 32 GB Mac for Kev-4B and Kev-9B. The kev-family release table instead names GPU memory: 4 GB for Kev-0.8B, 12 GB for Kev-4B and 24 GB for Kev-9B, with Apple Silicon noted in the same cell for the 0.8B and 4B rows but not for the 9B. Those are not the same floor. Unified memory is the single pool that Apple Silicon shares between the CPU, the GPU and everything you have open. As a result, a Mac has to fit the OS and your editor into the same RAM a CUDA card keeps to itself.

Unified memoryKev-0.8BKev-4BKev-9BKev-27B
8 GBDocumented: any Apple Silicon MacNoNoNo
16 GBDocumented, and slow: about 0.33 s per five-question request on an M5Not documented on a Mac; 12 GB is a CUDA figureNoNo
24 GBDocumentedThird-party measurement exists (see below), author says 32 GBNot documented on a Mac; 24 GB is a CUDA figureNo
32 GBDocumentedDocumented by the author, no memory reading publishedDocumented by the author, no memory reading publishedNo
64 GBDocumentedDocumentedDocumentedBorderline, per the model card
96-128 GBDocumentedDocumentedDocumentedExpected, never measured at 27B

Read the table as what the project claims, not as what has been verified. For Kev-0.8B the claim is unusually strong, because the README's "Runs on" column says any Apple Silicon Mac, and the 0.8B card's own Mac caveat is latency rather than memory. That caveat is roughly 0.33 s for a five-question request in bf16 on an M5, against 0.12 s for the older 0.6B. For Kev-4B and Kev-9B, by contrast, the Apple Silicon floor is not stated anywhere beyond that 32 GB cell. Notably, the 9B card states no Mac memory figure at all. From our analysis of all four cards, the memory floor simply does not exist as a published number for the two middle sizes.

What a first run actually downloads

This is the part that catches people who expect a 4B quant. The released tarballs carry the adapter, the pointer head, the tokenizer and the evaluation output, not the model. On first load the server pulls the frozen Qwen base from the Hub, and the adapters pin exact base revisions: Qwen/Qwen3.5-0.8B-Base at dc7cdfe2, Qwen/Qwen3.5-4B-Base at 1001bb4d and Qwen/Qwen3.5-9B-Base at 68c46c4b, as recorded in the 0.8B, 4B and 9B cards. Therefore the disk cost is the base, not the Kev file.

Consider the sizes, because this is where a budget surprise hides. The Qwen3.5 base repos are the bulk. Specifically, the 0.8B base is under 2 GB of files, the 4B base about 9 GB and the 9B base about 19 GB in bf16. That is why a "4B adapter plus base" download is a different order of magnitude from a 4B 4-bit quant. The adapters are small by comparison, because they are rank-16: 11.3M trainable parameters for Kev-0.8B, 33.8M for Kev-4B and 45.4M for Kev-9B, with the pointer head on top of each. Kev-27B has no adapter at all, which is the whole reason its card describes a 51.3 GB checkpoint.

The first hard Mac memory measurement is not from the author

A silver compact desktop workstation and a stack of solid-state drives lit by cyan-teal accent light on a dark shelf.

The strongest Apple Silicon evidence in the ecosystem is an unofficial build. RoderickQiu/kev-4b-mlx-8bit merges the Kev-4B LoRA into the Qwen3.5-4B base in fp32, quantises the result to 8 bits for MLX, and keeps the pointer head in fp32 including its fitted temperature. It is explicitly not an official Kev release. It was made for Qualm, a macOS app that uses Kev to tell a lecture from a feed.

Its table separates the two workloads that matter, and it was measured on an M5 Pro with 24 GB. For instance, the build path and the scoring path carry very different footprints:

Build it yourself (bf16)The 8-bit MLX file
Download9.0 GB (base plus adapter)4.5 GB
Bring-upabout 1 minute and a 15 GB peak to merge and quantise, a bf16 load peaks at 16 GBabout 4.8 GB
Memory while scoringabout 10 GBabout 6.6 GB

The 8-bit process reached 8.5 GB during the longest suite. That is the number to plan around for Kev-4B on a smaller Mac, because the bf16 path through Kev's own loader peaked at 16 GB just to load. As a result, a 16 GB Mac is out of room for Kev-4B in bf16, regardless of what the 12 GB CUDA cell suggests, while the quantised route fits a 24 GB machine with App memory to spare. Kev's own MLX path does have a published load behaviour for a Kev-4B written out as full bf16 weights: the README reports a load peaking at 8.4 GB for 8.4 GB of weights, against 15.9 GB for the adapter path. The extra copy is unavoidable while the LoRA is folded in.

Quality survives that conversion, on the development splits the third party actually scored: transfer-v4 0.817, decision-v7 0.873, hard-v1 0.786, devtools-v1 0.746 and documents-v1 0.896. Those sit against a local bf16 run of Kev's own MLX path on the same Mac. In particular, the same top answer appeared on 652 of 656 transfer-v4 questions and 1,262 of 1,264 decision-v7 questions. Those are development splits, so the locked test splits were not scored.

Three more community artifacts exist, and the pointer head is not the weak link in any of them. That corrects the usual assumption:

  • onnx-community/kev-4b-ONNX exports the graph for Transformers.js, with q4f16 and q4 variants and the pointer head kept in fp32. Loading it in llama.cpp will not work, and it is not a text generator.
  • FluidInference/kev-0.8b-coreml folds the 0.8B LoRA into the base in fp32 and ships Kev's pointer head with the Core ML packages, sized for states of 32-384 tokens.
  • mys/kev-4b-GGUF is GGUF-shaped but compiled by ggmlc, not llama.cpp, with the LoRA merged in fp32 before export and the sequence program inside the file. Notably, the repo states plainly that loading these files in llama.cpp will fail, which is the trap worth knowing before you download 4.02 GB of UD_Q4_K_M and reach for llama-cli.

None of these four is official. In addition, Ollama's registry has no kev or kev-4b tag today, and none of them is a served endpoint with a support promise behind it.

Kev-27B on a Mac is a promise, not a result

Kev-27B is where the Mac story stops being measurable. The model card describes a 51.3 GB bf16 checkpoint served in bf16 only, and gives the only measured memory numbers in the project: 65.5 GB resident on an H200 with a 17.6 s load from a warm cache, a 32k-token state peaking at 78.7 GB and a 64k-token state peaking at 87.1 GB. Therefore the README's own framing is "an 80 GB card", and the longest states need more than that.

The Apple Silicon path arrived on 2026-09-30, in a commit that makes the MLX backend load full-weight checkpoints as saved with nothing merged. The same card and the README both describe the result as an expectation: about 51 GB of weights plus working memory, so a 64 GB Mac is borderline and a 96-128 GB Mac should fit. Neither has been measured at 27B. The only real Mac run at this size was 27B version 1, which was still a LoRA adapter. An external contributor served it through MLX on a 128 GB M5 Max, at 52 GB steady and 97 GB at peak while the adapter was merged, with 0.849 accuracy on transfer-v4 development against the published 0.848. That is encouraging, and it is not the v2 checkpoint.

Treat the 96 GB figure as a derived estimate rather than a tested floor, and expect the working-memory margin to decide whether the server answers or swaps. If you want a 27B-class decision model on a Mac today, the honest position is that nobody has published a Kev-27B Mac run to copy.

The accuracy numbers belong to the author

The family's headline results are the author's measurements on the author's frozen suites, and they should be read that way. Kev-27B scores 0.851 development and 0.889 test on new sources with Brier 0.225 and 0.156, which is within a point of Jev's 0.857 development on the same new-source items. The author writes explicitly that this is not a controlled comparison, because Jev's training data is unknown. Meanwhile, Kev-4B and Kev-9B land within four points of Jev on the same items. Kev-0.8B is a sub-1B model that trails on purpose, which makes it the size choice rather than the accuracy choice.

Two caveats carry over to any workload. Calibration is a single fitted temperature, so the share of decisions you can automate at a 5% error budget is 0.14 for Kev-0.8B, 0.52 to 0.69 for Kev-4B, 9B and 27B, and 0.70 for Jev. Therefore, check the threshold on your own data. Knowledge questions, on the other hand, are set by the base model rather than by Kev, which is the same gap we described for CLM-8B in this category.

Who should download it this week

  • 8-16 GB Mac: Kev-0.8B is the documented and official choice, and it is the only size the project claims for any Apple Silicon Mac. Expect hundreds of milliseconds per five-question request, not tens.
  • 24-32 GB Mac: Kev-4B is the recommended default in the README and fits your machine, yet only the unofficial 8-bit MLX file has a published Mac memory figure, at about 6.6 GB while scoring and 8.5 GB peak on the longest suite. For example, a 24 GB Mac, such as the M5 Pro the third party tested, runs that build with room to spare, while the official bf16 load peaked at 16 GB in that same test.
  • 32 GB Mac and up, want the most accurate Kev: Kev-9B is the step up, with no published Mac memory reading to plan against. Verify on your own machine before you commit to it.
  • 64-128 GB Mac: Kev-27B is untested at 27B on any Mac. Wait for someone to publish a run, or run version 1's adapter path if you want to see the shape of the cost.
  • Anything older than an M-series Mac: out of scope. MLX is Apple Silicon only, and Kev-27B's card drops anything smaller than an 80 GB-class GPU or a Mac with about 96 GB of memory.

For the general sizing logic behind those numbers, our guide to how much VRAM an LLM needs covers the working-memory margin that matters more than the weight size. Every figure above was checked against the source linked beside it as part of our editorial process, and ModelFit runs no benchmarks of its own: the recommendation engine computes what fits. About ModelFit explains how that works.

FAQ

Does Kev replace Jev without code changes?

For the TypeSafe Python SDK, yes by design: the SDK is included with uv sync --extra serve and runs against a local Kev server unchanged. What changes is the model behind the API, not your call sites. Whether the answers are good enough is a separate question, and the author's own caveat is that the Jev comparison is not controlled.

Is there an official Ollama tag or GGUF for Kev?

No. In our experience, the assumption that every model ends up on Ollama breaks on decision models, and the checks agree: both ollama.com/library/kev and ollama.com/library/kev-4b return 404 at the time of writing, and the GGUF files that exist are ggmlc builds that the repository says llama.cpp will fail to load.

Why does Kev-4B need 32 GB when the release table says 12 GB?

Because those are two different resources. The 12 GB is GPU video memory for the CUDA path, where weights and buffers live alone. On a Mac the same weights come out of unified memory shared with macOS and everything else you have open, which is why the README's Mac column is stricter and why the one measured bf16 load needed 16 GB of peak.

Can I fine-tune a Kev on my Mac?

The author's path is PyTorch with --device cuda, and the training recipe assumes an H100-class run, at about 20 minutes for the first Kev-0.8B stage. A small adapter fine-tune does fit a 4 GB GPU in bf16 with batch 1. However, the published recipe is not a Mac recipe, and the card warns that two training jobs on the same Apple GPU are much slower.

What is the difference between choice, noul and score?

They are the three question types in one request, all scored in the same forward pass. choice picks from a named option set and returns a probability per option, noul returns a yes/no probability, and score returns a position on an ordered rubric with a confidence. Nothing is generated as text, so you never parse a label out of a completion.

Sources

What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

The weekly local-AI refresh

New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.

By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.

Have questions? Reach out on X/Twitter