By Peter · ModelFit · 2026-10-07

AREX-2 on a Mac: 17.20 GB at Q4 and a 16 GiB KV Cache

Illuminated AREX logo on a black panel beside a laptop and a compact desktop, with a long row of memory modules glowing behind them

BAAI published AREX-2, a 27B multimodal agent model, on 29 September 2026 under Apache-2.0. The community had GGUF quants out the next day. That turns the usual "will it ship" question into a sizing question. Specifically, the weights fit on a mid-range Mac, but the headline 262,144-token context does not come free. Below is what each quant weighs, which Mac tier takes it, and how much memory the long context adds once you actually fill it.

What AREX-2 is

AREX-2 is a 27B dense model from the Beijing Academy of Artificial Intelligence, built on Qwen3.8-27B. In fact, its card describes the job as long-horizon self-improvement. The model proposes a solution, measures it against verifiable feedback, reflects, and revises across multiple test-time rounds. It was trained on machine-learning and algorithmic-programming tasks, where a score or a log tells the model whether it improved. According to the card, the same behavior transfers to deep research without new search trajectories.

The official model card lists three things that matter for a Mac:

  • Architecture: dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

The paper is at arXiv 2609.38288, submitted 29 September 2026. Meanwhile, the code repository the card points to is github.com/VectorSpaceLab/AREX-2. It was opened during this check and holds the AREX-2 evaluation and data scripts, not the earlier AREX project.

Licence is worth stating plainly. The checkpoint is Apache-2.0, but it is a finetune of Qwen3.8-27B, and the card itself says to follow the terms and notices for the Qwen base model. Therefore, if you download AREX-2, you are also relying on Qwen base weights under Qwen's terms.

Which Mac runs AREX-2

Quantizations exist already. bartowski's GGUF repo lists the standard ladder, and the sizes below are the decimal-GB figures from its own file table.

QuantFile sizeNotes
Q2_K10.58 GBVery low quality but loadable.
Q4_K_M17.20 GBThe default pick.
Q5_K_M20.68 GBHigher quality, larger.
Q8_028.67 GBNear-lossless, rarely needed.

Matching those against Mac memory is the part the file table does not answer. Unified memory is the single pool Apple Silicon shares between CPU and GPU, and macOS caps how much of it the GPU may wire by default. As a result, the budget available to llama.cpp sits below the marketed RAM figure. With that in mind:

Mac unified memoryQuant that fitsPractical contextNotes
16 GBQ2_K (10.58 GB)Short, around 8KThe floor, not a daily driver.
24 GBQ3_K_M (13.17 GB)32KQ4_K_M fits the weights but leaves no room for KV.
32 GBQ4_K_M (17.20 GB)32K to 64KThe first tier where the default quant is comfortable.
36 GB to 48 GBQ4_K_M or Q5_K_M (20.68 GB)128KRoom for a long-context KV cache.
64 GBQ8_0 (28.67 GB) or Q4_K_M262K with Q4_K_MHeadroom for KV and the vision projector.
96 GB and upQ8_0 (28.67 GB)262KQ8_0 plus a full-length KV cache fits.

The short version: Q4_K_M at 17.20 GB is a 32 GB Mac story once you allow any real context. Notably, anything smaller pushes you to Q3 or Q2. If you are still choosing hardware, the sizing logic in our VRAM guide and the Mac mini M5 Pro 64GB breakdown apply here unchanged.

The 262K context is the real memory story

Close-up of memory chips on a dark circuit board, a few lit in cyan while the rest stay in shadow

The weight file is the number everyone quotes. For AREX-2, however, the KV cache at full context is nearly as large, and that is the figure worth planning around. KV cache is the per-token memory a model keeps for the keys and values it has already attended to.

AREX-2 does not use plain full attention on every layer. Full attention is the layer type whose cache grows with context length, while linear attention is the variant that keeps a fixed-size recurrent state instead. Specifically, the published config.json sets full_attention_interval to 4, and its layer_types array lists three linear_attention layers followed by one full_attention layer, repeated across 64 hidden layers. That gives 16 full-attention layers and 48 linear-attention layers. Because the linear layers keep a recurrent state whose size does not grow with context, only the 16 full-attention layers carry a KV cache that scales with length.

Do the arithmetic with the same config. Each full-attention layer has 4 key-value heads at a head dimension of 256. Thus, at f16, each layer stores 2 (key and value) x 4 x 256 x 2 = 4096 bytes per token. Across 16 layers that is 65536 bytes, or 64 KiB, per token.

  • 32,768 tokens: about 2 GiB of KV
  • 131,072 tokens: about 8 GiB of KV
  • 262,144 tokens: about 16 GiB of KV

Read that against the table above. At the full 262,144-token context, the KV cache alone is about 16 GiB, roughly the size of the Q4_K_M weight file. On a 32 GB Mac, consequently, that leaves little for anything else; on a 64 GB Mac it is comfortable. This is arithmetic from the published config, not a measurement, and it assumes f16 cache with no quantized-KV tricks. Quantizing the cache to q8_0 or q4_0 shrinks it further, at a small quality cost.

This is the part of the spec that changes the plan: only 16 of the 64 layers carry a cache that scales with length. The marketed RAM figure is the wrong number to plan against. The wired-GPU budget is the one that bites.

It could have been worse. Because only 16 of the 64 layers carry a KV cache, the full-context figure is about 16 GiB. A fully dense attention stack at this head count, in contrast, would imply a number roughly four times larger.

Does the vision projector work in llama.cpp

bartowski ships two projector files alongside the quants: mmproj-BAAI_AREX-2-f16.gguf at 927,607,264 bytes and mmproj-BAAI_AREX-2-bf16.gguf at 931,146,208 bytes. Their presence is not the same as a working pipeline.

As of this check, no public report confirms that AREX-2's multimodal input runs through llama.cpp on Apple Silicon, or that the model keeps its agent behavior once quantized to Q4. For example, the model card's own inference example uses Transformers with AutoModelForMultimodalLM, not llama.cpp. Vision support for this architecture family is still settling in llama.cpp. In other words, a projector file being published does not establish that image input works. Treat AREX-2's vision path as unverified on a Mac until someone reports it running. The text path is the safer expectation.

The benchmark numbers are BAAI's own

The model card reports AREX-2 at 70.7 on Frontier-CS, 81.8 on MLE-Lite, 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA. Those results follow the protocols in the AREX-2 paper. In addition, they are BAAI's own measurements. Notably, no third party has reproduced them, and a 27B model beating frontier closed models on the card's own tables deserves exactly that caveat. The card is future reach and agentic tasks, not everyday chat. If you want a general Mac assistant first, our 30B agent roundup and the Qwen3.8-27B finetune notes are closer to that question.

We publish no speed numbers for AREX-2. There is no Apple Silicon tokens-per-second measurement for this model anywhere, official or community, and ModelFit runs no benchmarks of its own. Therefore, if you want a throughput figure, measure it locally rather than trusting a number from a different machine.

FAQ

Can a 16 GB Mac run AREX-2?

Only at the bottom of the ladder. In fact, Q2_K is 10.58 GB, which leaves almost nothing for KV cache on a 16 GB Mac once macOS takes its share. As a result, expect a short context and mediocre output. Q4_K_M at 17.20 GB does not fit.

Is there an MLX version of AREX-2?

Not as of this check. The Hugging Face search for AREX-2 returns GGUF, Transformers, FP8 and INT8 and INT4 conversions, but no MLX conversion, official or community. A similarly named community model exists, but it is not AREX-2.

How much memory does the full 262,144-token context need?

About 16 GiB of KV cache at f16, from the published config. That assumes the hybrid attention layout is honored. Alternatively, a quantized KV cache reduces it.

Does the vision projector work in llama.cpp?

Unverified. The f16 and bf16 projector files are published, but no public report confirms image input through llama.cpp on Apple Silicon, and the official example uses Transformers instead.

Do I need the mmproj file at all?

Only for image input. The text path runs from the quant alone. If you load the projector and multimodal input fails, drop --mmproj and run text.

Sources

What hardware runs this?

Match this model to a machine that can run it: by RAM tier for Apple Silicon, or by VRAM for an NVIDIA GPU.

See how this changes your recommendation
Run the wizard

Have questions? Reach out on X/Twitter