Laya is not a chatbot and it never writes a token. It reads a state (an email, a ticket, a JSON blob) plus typed questions, and it returns calibrated probabilities in one forward pass. Convai Innovations released it under Apache-2.0 on 18 September 2026, and a community MLX port reported 13.42 ms for a short question on an M3 Max a day later. This article separates the vendor numbers from the author numbers, and it explains what "13ms" measures. In addition, it answers the only question that matters on a Mac: does it run on yours, and should you load it through MLX, through the Neural Engine, or not at all?
What Laya actually is
Laya is a non-autoregressive System 1 decision model, which means it answers in a single forward pass instead of generating one token at a time. A fully fine-tuned ModernBERT-large encoder (395M parameters, bidirectional) forms the backbone, and on top of it sits a decision head. ModernBERT is a bidirectional encoder architecture, and this model uses the large variant as its backbone. A decision head is the small stack that turns the encoder's output into typed answers: two transformer layers, an option-marker scorer, and an act/escalate head. Total parameters: 421,293,830, which the model card rounds to 421M. It includes a multilingual variant and a typed-decisions variant, which is why the family name covers three checkpoints rather than one.
The output is constrained, not generated. Every question declares a type:
choice: probabilities over named options, plus a calibrated confidence.score: a probability distribution over ordered rubric levels, plus the expected level.noul: a calibrated P(true) between 0.0 and 1.0 for a yes/no proposition.
There is no token-by-token decoding, no JSON for you to parse, and no hallucinated free text. That is the entire architectural bet: keep the model small enough to run on commodity hardware, and make its answers measurable.
However, the trade-off is a hard input budget. On the English convaiinnovations/laya checkpoint the limit is 512 tokens per question, and that budget covers the instructions, the options and the state together. Consequently, a long email will be truncated. If you need more room, the multilingual checkpoint ships a 1,024-token limit and can be pushed to 8,192 with max_len=8192, though the card reports accuracy degrading beyond roughly 4,000 tokens of state. The typed-decisions checkpoint also sits at 1,024. Check the JV System One prior art post for the category context this fits into.
The numbers, and who measured them
This is where most coverage goes wrong, so read the attribution before the latency. The 38.4 ms and 156.0 ms figures come from the Convai Innovations model card. The 13.42 ms and 4.98 ms figures, where they exist at all, come from port authors. Importantly, none of these are measured by modelfit. Here is the split, honestly labelled:
| Metric | Value | Measured by | Notes |
|---|---|---|---|
| One question, P50 | 38.4 ms | Convai model card | P95 42.1 ms |
| Ten batched questions | 156.0 ms | Convai model card | P95 158.4 ms |
| Fifty batched questions | 721.4 ms | Convai model card | parallel mini-batching |
| One short question, P50 | 13.42 ms | laya-mlx port author | M3 Max, FP16, 421M checkpoint |
| One short question, P50 | 7.39 ms | laya-mlx port author | M3 Max, multilingual 322M |
| Fifty-question throughput | 146.8 q/s | laya-mlx port author | batch_size 64 |
| Peak MLX allocation | 943.6 MiB | laya-mlx port author | one short question, 421M |
The model card also reports 83.8 percent in-task macro accuracy (macro accuracy 0.838) against 67.8 percent for TypeSafe Jev, on the card's own benchmark. Treat cross-vendor accuracy tables as indicative, not apples to apples. For instance, prompts and sample sizes differ, and Convai states plainly that Jev figures are third-party published, not re-measured.
The one independent measurement in circulation is a Sperix Labs campaign on an M4 Max, reporting 6.6 to 19.6 ms per question at 50-token states over 100 repeats. It is cited by stackness.dev and was not read at source for this article, so it is "cited, not verified". The laya-mlx benchmarks include a separate sustained-load test, also cited by stackness.dev and also not read at source. If you need a number you can defend in a meeting, measure it on your own machine.
That labelling is deliberate. Our editorial rule is simple: no figure ships without a named source, and we would rather write not verified than guess. If you spot an error in that table, contact us and we will correct it in the article, because a wrong latency number is worse than no number at all. Who we are is on the about page.
Should you run Laya through MLX, ONNX on the CPU, or the Neural Engine?
Three runtimes are in play, and they are not equal on a Mac.
MLX is Apple's array framework for Apple Silicon, and via the laya-mlx port it is the only path that puts the 421M encoder on the Apple GPU. That port is a Python library, not a server: you pip install laya-mlx and call predict() in-process. This is the route the 13.42 ms figure belongs to.
ONNX on the CPU, via Ollaya, is the server route. Ollaya is the local runtime for decision models (ollaya-dev/ollaya), and it runs them through ONNX Runtime on the CPU, and through CUDA on NVIDIA GPUs. There is no MLX, Metal or Neural Engine backend for Laya in Ollaya today. On a Mac, therefore, Ollaya serves Laya on the CPU: correct and convenient, but not the same speed class as the MLX port. If a post tells you Ollaya accelerates Laya on Apple Silicon, that is not what its documentation says.
In contrast, the Neural Engine route, via a claimed Core ML build, cannot be verified. The proposal for this article referenced a Core ML README quoting 4.98 ms P50 on an M3 Max Neural Engine, attributed to the same port author. We found no such README at source, so that figure is "not verified" and does not appear in the table above. If a Core ML variant does land, it would be worth its own benchmark.
One routing warning that changes the recommendation: laya-mlx is an independent MLX port by mizorewww (mizorewww/laya-mlx), not an official Convai Innovations release. It retains the original weights, question formatting, calibration and output schema. Notably, its validation matched the upstream selected answer on 63 of 63 questions in both FP32 and FP16. However, model quality and calibration limits remain those of the upstream checkpoint, and the port says so in its own README.
Can your Mac run it? Memory by tier
The 421M checkpoint is genuinely small. We ran that arithmetic ourselves: 421,293,827 weights at 2 bytes each is about 803 MiB, a derived figure from the published parameter count rather than a measurement. The port author measured a peak MLX allocation of 943.6 MiB for one short question. In particular, that 943.6 MiB is the peak MLX allocation, not total process memory and not a loaded-machine figure, so do not convert it into a "fits in X GB" claim. What follows is guidance, not a measured footprint:
| Mac unified memory | Laya 421M via laya-mlx | Verdict |
|---|---|---|
| 8 GB | Runs, little headroom | Weights plus runtime fit; a browser and an IDE alongside it do not |
| 16 GB | Comfortable | Room for the encoder and normal work simultaneously |
| 24-32 GB | Comfortable, room for a second checkpoint | Router can keep English and multilingual resident |
| 64 GB and up | No constraint | The model is not the bottleneck; your batch size is |
Similarly, the multilingual 322M checkpoint is smaller still, at a measured 687.6 MiB peak MLX allocation. For example, a 16 GB machine sits comfortably per that tier table, while 8 GB leaves little headroom. For model-vs-machine sizing in general, see how much VRAM for LLMs and the Splash engine Apple Silicon RAM breakdown.
There is one hard eligibility line worth stating. Specifically, laya-mlx requires Python 3.11 or later and macOS 14 or later. The author measured on macOS 27.2 with Python 3.12.13 and MLX 0.32.2, and noted that older supported macOS versions were not tested on that machine. An Intel Mac is out; this is Apple Silicon only.
MLX versus the Ollama-style runtime
Ollaya is the "Ollama for decision models" layer. It pulls models by name and serves them from a local daemon on port 11435. In addition, it speaks a wire format it calls identical to TypeSafe's /v1/systemone, so the official TypeSafe Python SDK works unchanged once you point it at localhost:
export TYPESAFE_BASE_URL=http://localhost:11435
That compatibility is the real product. If you already have code written against the TypeSafe SDK, switching the base URL is the whole migration.
Thus the split between the two runtimes maps cleanly onto their job:
- Want speed on a Mac and can live inside a Python process? Use
laya-mlx. - Want a daemon, a model registry and a drop-in TypeSafe endpoint? Use Ollaya, and accept CPU speed on macOS.
- Want low-level control over ONNX graphs and a GPU on Linux? Ollaya with CUDA.
Do not expect to load Laya in Ollama, LM Studio or llama.cpp. Laya is a ModernBERT encoder with a custom decision head and ships no GGUF build, so none of those tools can load it. In fact, no Ollama or LM Studio tag exists for the Laya decision models today, and the only thing named Laya in those registries is an unrelated notification app. That story is covered in more depth in our MLX versus Ollama on Mac comparison.
The calibration caveat you should not skip
A decision model lives or dies on whether its confidence means anything, and here the news is mixed.
Following upstream v0.3.5, laya-mlx clamps fitted calibration temperatures to the range 0.5 to 5.0 before use. The shipped choice:11+ bucket carries a temperature of 0.1006, and that raw value would sharpen the logits by roughly 10x, reporting a coin flip as near-certainty. The port clamps it and emits a RuntimeWarning at load naming every clamped bucket. Meanwhile, the raw values remain available as temperature_raw and temperature_by_options_raw.
The practical consequence: if your workload uses that bucket, treat the reported confidence as uncalibrated until you measure it yourself. The port's guidance is explicit on this point. In other words, the same warning applies across runtimes, which is why Stackness argues you should pin the runtime version, not just the model version, when you threshold on a probability.
The upside is that the decision head is calibrated by design. Training used RLCD (Reinforcement Learning for Calibrated Decisions), where the reward is a strictly proper scoring rule, so maximum expected reward is reached only when the reported probabilities are honest. On the typed-decisions checkpoint, expected calibration error ranges from 0.192 on noul to 0.255 on choice over 2,000 decisions. That said, low ECE is the point of the architecture, and you should still verify it on your own distribution before you automate on it.
FAQ
Does Laya generate text?
No. It is a non-autoregressive encoder with a decision head. It returns a selected option, a rubric score and calibrated probabilities, never a generated token or a JSON string to parse.
Can I run Laya in Ollama or LM Studio?
No. Laya has no GGUF build, so llama.cpp, Ollama and LM Studio cannot load it. Instead, use laya-mlx for in-process MLX inference, or Ollaya for a local server.
How much RAM does Laya need on a Mac?
At FP16 the 421M weights are about 803 MiB, and the port author measured a peak MLX allocation of 943.6 MiB for one short question. Total process memory was not measured, so, as a result, the safe answer is this: any 16 GB Apple Silicon Mac handles it comfortably, and 8 GB works with little headroom.
Is the 13.42 ms figure a modelfit measurement?
No. It is the laya-mlx port author's measurement on an M3 Max, 40 GPU cores, 128 GiB, FP16, one short question, model loading excluded. The 38.4 ms figure is Convai's. We measured neither.
Does Ollaya use the GPU on a Mac?
Not for Laya as documented. Ollaya runs decision models through ONNX Runtime on the CPU, and CUDA on NVIDIA GPUs. Notably, on Apple Silicon that means the CPU path, while the GPU path for Laya is the separate laya-mlx library.
What this means for a Mac local-AI stack
Laya fills a slot that chat models cannot: fast, typed, calibrated decisions with zero generation and a tiny footprint. On a 16 GB Apple Silicon Mac, laya-mlx runs it in-process at roughly 13 ms per short question by the author's own measurement. If you need a daemon that speaks TypeSafe, Ollaya gives you that on the CPU. What you cannot do is smuggle a ModernBERT decision encoder into your existing GGUF toolchain: the formats do not overlap, and that is a design choice, not a gap waiting to be patched. For example, a triage step that decides whether a support ticket needs a human can call laya-mlx in-process and keep every decision on the machine.
Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.
The weekly local-AI refresh
New open-weight models, real Apple Silicon benchmarks, and the one model worth running on your Mac this week. Free, one email a week, unsubscribe anytime.
By subscribing you agree to our Privacy Policy and to receive the weekly email. Unsubscribe anytime.
Have questions? Reach out on X/Twitter