LFM2 24B-A2B Instruct quants compared

27 GGUF builds by real file size, probed from bartowski/LiquidAI_LFM2-24B-A2B-GGUF on Hugging Face (2026-09-02). 24B params, 2B active.

Download LFM2 24B-A2B Instruct Q4_K_M (13.44 GB) — it fits 24 GB of memory with 16k context. With Ollama: ollama run lfm2:24b-a2b

Every LFM2 24B-A2B Instruct quant by real file size

QuantWeights+ KV (16k)TotalFits comfortably inQuality
BF1644.42 GB4.0 GB48.4 GB64 GBFull precision (lossless)
Q8_023.61 GB4.0 GB27.6 GB32 GBNear-lossless
Q6_K_L18.27 GB4.0 GB22.3 GB32 GBExcellent
Q6_K18.24 GB4.0 GB22.2 GB32 GBExcellent
Q5_K_L15.8 GB4.0 GB19.8 GB24 GBVery high
Q5_K_M15.77 GB4.0 GB19.8 GB24 GBVery high
Q5_K_S15.31 GB4.0 GB19.3 GB24 GBVery high
Q4_K_M *13.44 GB4.0 GB17.4 GB24 GBHigh — the default pick
Q4_113.93 GB4.0 GB17.9 GB24 GBHigh
Q4_K_L13.47 GB4.0 GB17.5 GB24 GBHigh
Q4_K_S12.99 GB4.0 GB17.0 GB24 GBHigh
Q4_012.77 GB4.0 GB16.8 GB24 GBHigh
IQ4_NL12.56 GB4.0 GB16.6 GB24 GBHigh
IQ4_XS11.87 GB4.0 GB15.9 GB24 GBHigh
Q3_K_XL10.53 GB4.0 GB14.5 GB24 GBAcceptable — visible loss
Q3_K_L10.5 GB4.0 GB14.5 GB24 GBAcceptable — visible loss
Q3_K_M10.1 GB4.0 GB14.1 GB16 GBAcceptable — visible loss
IQ3_M10.09 GB4.0 GB14.1 GB16 GBAcceptable — visible loss
Q3_K_S9.64 GB4.0 GB13.6 GB16 GBAcceptable — visible loss
IQ3_XS9.11 GB4.0 GB13.1 GB16 GBAcceptable — visible loss
IQ3_XXS8.75 GB4.0 GB12.8 GB16 GBAcceptable — visible loss
Q2_K_L7.78 GB4.0 GB11.8 GB16 GBExperimental — not ranked — never recommended
Q2_K7.75 GB4.0 GB11.8 GB16 GBExperimental — not ranked — never recommended
IQ2_M7.21 GB4.0 GB11.2 GB16 GBExperimental — not ranked — never recommended
IQ2_S6.45 GB4.0 GB10.4 GB12 GBExperimental — not ranked — never recommended
IQ2_XS6.17 GB4.0 GB10.2 GB12 GBExperimental — not ranked — never recommended
IQ2_XXS5.35 GB4.0 GB9.3 GB12 GBExperimental — not ranked — never recommended

* default pick. Weights = real GGUF file sizes from bartowski/LiquidAI_LFM2-24B-A2B-GGUF (probed 2026-09-02). KV = fp16 estimate; a q8_0 cache roughly halves it. "Comfortable" = weights + KV within 90% of memory.

Best LFM2 24B-A2B Instruct quant by memory

MemoryRecommended quantTotal (16k ctx)
16 GBQ3_K_M14.1 GB
24 GBQ5_K_M19.8 GB
32 GBQ8_027.6 GB
64 GBBF1648.4 GB

Why we don't rank LFM2 24B-A2B Instruct's 2-bit quants

Quants at 2 bits per weight or below (Q2_K, IQ2, IQ1, TQ1) cut file size by roughly half versus Q4, but the quality collapse is steep and non-linear: perplexity spikes, instruction-following degrades, and hallucinations rise. A model that answers faster but wrong is not a smaller model — it is a worse one. ModelFit lists these builds for completeness but never ranks or recommends them.

Frequently asked questions

What is the best quantization of LFM2 24B-A2B Instruct?

Q4_K_M is the default pick: 13.44 GB of weights, high — the default pick quality, fitting comfortably in 24 GB of memory (weights + 16k context KV-cache). Go Q6_K or Q8_0 if you have headroom.

How much memory does LFM2 24B-A2B Instruct need?

At Q4_K_M, LFM2 24B-A2B Instruct needs 13.44 GB for the weights plus ~4.0 GB of KV-cache at 16k context — about 17.4 GB total, so a 24 GB card or Mac (90% usable budget) runs it comfortably.

Should I use a Q2_K or IQ2 quant of LFM2 24B-A2B Instruct?

No. LFM2 24B-A2B Instruct at 2 bits per weight is a visibly worse model — quality collapse at that bitrate is steep, not gradual. If only a 2-bit build fits your memory, run a smaller model at Q4_K_M instead. ModelFit lists these builds but never recommends them.

Cite this page

ModelFit: LFM2 24B-A2B Instruct quantization comparison (real GGUF file sizes).
https://modelfit.io/quant-compare/lfm2-24b-a2b/ (data probed 2026-09-02, CC BY 4.0).