LFM2.5 8B-A1B

LFM2.5 8B-A1B loads in 5.5 GB at Q4_K_M and is comfortable from 12 GB of memory, Macs included from the Mac Mini M6 16GB. Below: the full memory math, the cheapest card that runs it well, and the fastest machine we track.

PARAMETERS
8.3B (1.5B active)
FORMAT
Q4_K_M
MIN MEMORY
8 GB
BEST FOR
On-device agents, tool calling, multilingual chat

Memory math

Weights (Q4_K_M)
5.5 GB
+ KV cache (16k)
~2.0 GB
Total at 16k
~7.5 GB
Comfortable from
12 GB

KV = fp16 estimate (q8_0 cache roughly halves it). "Comfortable" = weights + KV within the engine's tiered budget (~70-85% of memory).

Hardware snapshot

Cheapest GPU
~148 tok/s est. — ~$550 used
Fastest
~291 tok/s est.
Runs on a Mac?
~64 tok/s est.

More LFM2 models

Frequently asked questions

How much memory does LFM2.5 8B-A1B need?

5.5 GB for the Q4_K_M weights, plus ~2.0 GB of KV-cache at 16k context — about 7.5 GB total. Comfortable from 12 GB of VRAM or unified memory.

Does LFM2.5 8B-A1B run on a Mac?

Yes — from the Mac Mini M6 16GB (~64 tok/s est.). Unified memory means the RAM budget is the only limit.

What is the cheapest GPU for LFM2.5 8B-A1B?

The AMD Radeon RX 7900 XT is the cheapest tracked card that runs LFM2.5 8B-A1B comfortably — ~148 tok/s est. at ~$550 used.

Cite this page

ModelFit: LFM2.5 8B-A1B — specs, memory math and hardware verdicts.
https://modelfit.io/models/lfm2.5-8b-a1b/ (dataset updated 2026-09-03, CC BY 4.0).