Local LLM Compatibility Dataset

107 models by params, quantization, and memory load. Which run locally on Apple Silicon and NVIDIA GPUs. Free under CC BY 4.0. Updated 2026-07-25.

This is ModelFit's open compatibility dataset: every model the site tracks, with the parameter size, quantization, minimum RAM, and estimated memory load used to decide what runs locally. Memory figures are system/unified RAM; the same budget math applies to GPU VRAM (per-card breakdowns live in the GPU guides). Reuse it freely with attribution (CC BY 4.0). Credit ModelFit (modelfit.io). Machine-readable version: /api/dataset/, or get the CSV + JSON on GitHub.

Prefer the terminal? The same dataset and engine power the open-source CLI: npx @wecko-ai/modelfit names the best local model for the machine it runs on. Get it on npm or GitHub.

107 of 107 models
FamilyQuantLocalollama
Claude 3.5 SonnetClaudeUndisclosedAPI0 GBCloud
Claude 3.7 SonnetClaudeUndisclosedAPI0 GBCloud
Claude 3 OpusClaudeUndisclosedAPI0 GBCloud
Claude 4 OpusClaudeUndisclosedAPI0 GBCloud
Claude 4 SonnetClaudeUndisclosedAPI0 GBCloud
Claude Opus 4.7ClaudeUndisclosedAPI0 GBCloud
Claude Opus 4.8ClaudeUndisclosedAPI0 GBCloud
Claude Fable 5ClaudeUndisclosedAPI0 GBCloud
DeepSeek-R1 Distill Qwen 7BDeepSeek7BQ4_K_M8 GB~5.5 GBYesdeepseek-r1:7b
DeepSeek-R1 Distill Qwen 14BDeepSeek14BQ4_K_M16 GB~11 GBYesdeepseek-r1:14b
DeepSeek-R1 Distill Llama 70BDeepSeek70BQ4_K_M64 GB~42 GBYesdeepseek-r1:70b
DeepSeek V4 FlashDeepSeek284BAPI0 GBOpen (llama.cpp)
DeepSeek-R1 671BDeepSeek671BQ4_K_M512 GB~380 GBYesdeepseek-r1:671b-q4_K_M
DeepSeek-V3DeepSeek671BAPI0 GBOpen, no fit
DeepSeek-V3-0324DeepSeek671BAPI0 GBOpen, no fit
DeepSeek-R1DeepSeek671BAPI0 GBOpen, no fit
DeepSeek V4 ProDeepSeek1600BAPI0 GBOpen, no fit
Gemma 3 1B InstructGemma1BQ4_K_M2 GB~1 GBYesgemma3:1b
Gemma 2 2B InstructGemma2BQ4_K_M4 GB~1.8 GBYesgemma2:2b-instruct-q4_K_M
Gemma 4 E2BGemma2.3BQ4_K_M4 GB~2.3 GBYesgemma4:e2b
Gemma 3 4B InstructGemma4BQ4_K_M6 GB~3.5 GBYesgemma3:4b
Gemma 4 E4BGemma4.5BQ4_K_M6 GB~4 GBYesgemma4:e4b
Gemma 2 9B InstructGemma9BQ4_K_M12 GB~7 GBYesgemma2:9b-instruct-q4_K_M
Gemma 3 12B InstructGemma12BQ4_K_M16 GB~9.5 GBYesgemma3:12b
Gemma 4 12BGemma12BQ4_K_M12 GB~8 GBYesgemma4:12b
Gemma 4 12B (Q8)Gemma12BQ8_024 GB~12.8 GBYesgemma4:12b-it-q8_0
Gemma 4 26B-A4BGemma26BQ4_K_M24 GB~16 GBYesgemma4:26b
Gemma 4 26B-A4B (Q8)Gemma26BQ8_048 GB~28.1 GBYesgemma4:26b-a4b-it-q8_0
Gemma 2 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma2:27b-instruct-q4_K_M
Gemma 3 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma3:27b
Gemma 4 31BGemma31BQ4_K_M32 GB~20 GBYesgemma4:31b
Gemini 2.5 ProGoogleUndisclosedAPI0 GBCloud
Gemini 2.5 FlashGoogleUndisclosedAPI0 GBCloud
Gemini 3.1 ProGoogleUndisclosedAPI0 GBCloud
GPT-OSS 20BGPT-OSS21BMXFP424 GB~13.8 GBYesgpt-oss:20b
GPT-OSS 120BGPT-OSS117BMXFP496 GB~65.4 GBYesgpt-oss:120b
Granite 4.1 3B InstructGranite3BQ4_K_M4 GB~2 GBYesgranite4.1:3b
Granite 4.1 8B InstructGranite8BQ4_K_M8 GB~5.5 GBYesgranite4.1:8b
Kimi K2 InstructKimi1000BAPI0 GBOpen, no fit
Kimi K2.6Kimi1000BAPI0 GBOpen, no fit
Kimi K2.7-CodeKimi1000BAPI0 GBOpen, no fit
Poolside Laguna XS.2Laguna33BQ4_K_M36 GB~23 GBYeslaguna-xs.2:q4_K_M
LFM2.5 8B-A1BLFM28.3BQ4_K_M8 GB~5.5 GBYeslfm2.5:8b-a1b-q4_K_M
LFM2 24B-A2B InstructLFM224BQ4_K_M24 GB~14 GBYeslfm2:24b-a2b
Llama 3.2 1B InstructLlama1BQ4_K_M2 GB~1 GBYesllama3.2:1b-instruct-q4_K_M
Llama 3.2 3B InstructLlama3BQ4_K_M4 GB~2.5 GBYesllama3.2:3b-instruct-q4_K_M
Llama 3.1 8B InstructLlama8BQ4_K_M12 GB~6.5 GBYesllama3.1:8b-instruct-q4_K_M
Llama 3.1 8B Instruct (Q5)Llama8BQ5_K_M12 GB~8 GBYesllama3.1:8b-instruct-q5_K_M
Llama 3.1 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.1:70b-instruct-q4_K_M
Llama 3.3 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.3:70b-instruct-q4_K_M
Llama 3.3 70B Instruct (Q8)Llama70BQ8_096 GB~75 GBYesllama3.3:70b-instruct-q8_0
Llama 3.3 70B Instruct (Q6)Llama70BQ6_K96 GB~57.9 GBYesllama3.3:70b-instruct-q6_K
Llama 4 ScoutLlama109BQ4_K_M96 GB~67 GBYesllama4:scout
Llama 4 MaverickLlama400BQ4_K_M320 GB~245 GBYesllama4:maverick
Llama 3.1 405B InstructLlama405BQ4_K_M320 GB~243 GBYesllama3.1:405b-instruct-q4_K_M
Xiaomi MiMo-V2-FlashMiMo309BAPI0 GBOpen (llama.cpp)
MiniMax M3MiniMax428BAPI0 GBOpen, no fit
Mistral 7B InstructMistral7BQ4_K_M8 GB~5.5 GBYesmistral:7b-instruct-q4_K_M
Mistral Nemo 12BMistral12BQ4_K_M16 GB~9.5 GBYesmistral-nemo:12b
Mistral Small 22BMistral22BQ4_K_M32 GB~17 GBYesmistral-small:22b
Mistral Small 3.1Mistral24BQ4_K_M24 GB~15 GBYesmistral-small3.1:24b
Mixtral 8x7B InstructMistral46.7BQ4_K_M48 GB~30 GBYesmixtral:8x7b
Mistral Medium 3.5Mistral128BAPI0 GBOpen (llama.cpp)
NVIDIA Nemotron Cascade 2 30B-A3BNemotron30BQ6_K36 GB~24 GBYesnemotron-cascade-2:30b
NVIDIA Nemotron 3 UltraNemotron550BAPI0 GBOpen, no fit
Cohere North Mini CodeNorth30BQ4_K_M32 GB~19 GBYesnorth-mini-code-1.0:q4_K_M
GPT-4oOpenAIUndisclosedAPI0 GBCloud
GPT-4o miniOpenAIUndisclosedAPI0 GBCloud
GPT-5.5OpenAIUndisclosedAPI0 GBCloud
Phi-3 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi3:mini
Phi-4 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi4-mini:3.8b
Phi-3 Medium 14BPhi14BQ4_K_M16 GB~11 GBYesphi3:medium
Phi-4 14BPhi14BQ4_K_M24 GB~11.5 GBYesphi4:14b-q4_K_M
Qwen3.7-PlusQwenUndisclosedAPI0 GBCloud
Qwen2.5 0.5B InstructQwen0.5BQ4_K_M2 GB~0.8 GBYesqwen2.5:0.5b-instruct-q4_K_M
Qwen3.5 0.8B InstructQwen0.8BQ4_K_M2 GB~0.8 GBYesqwen3.5:0.8b
Qwen2.5 1.5B InstructQwen1.5BQ4_K_M4 GB~1.5 GBYesqwen2.5:1.5b-instruct-q4_K_M
Qwen3.5 2B InstructQwen2BQ4_K_M4 GB~1.8 GBYesqwen3.5:2b
Qwen2.5 3B InstructQwen3BQ4_K_M4 GB~2.5 GBYesqwen2.5:3b-instruct-q4_K_M
Qwen3.5 4B InstructQwen4BQ4_K_M6 GB~3.5 GBYesqwen3.5:4b
Qwen2.5 7B InstructQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5:7b-instruct-q4_K_M
Qwen2.5 Coder 7BQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5-coder:7b
Qwen3 8BQwen8BQ4_K_M12 GB~6.5 GBYesqwen3:8b-q4_K_M
Qwen3.5 9B InstructQwen9BQ4_K_M12 GB~7 GBYesqwen3.5:9b
Qwen3.5 9B Instruct (Q8)Qwen9BQ8_016 GB~10.7 GBYesqwen3.5:9b-q8_0
Qwen2.5 14B InstructQwen14BQ4_K_M16 GB~11 GBYesqwen2.5:14b-instruct-q4_K_M
Qwen2.5 Coder 14BQwen14BQ4_K_M16 GB~11 GBYesqwen2.5-coder:14b
Qwen3 14BQwen14BQ4_K_M16 GB~11 GBYesqwen3:14b-q4_K_M
Qwen3 14B (Q8)Qwen14BQ8_024 GB~15.9 GBYesqwen3:14b-q8_0
Qwen3.5 27B InstructQwen27BQ4_K_M24 GB~16 GBYesqwen3.5:27b
Qwen3.6 27BQwen27BQ4_K_M32 GB~18 GBYesqwen3.6:27b
Qwen3.6 27B (Q8)Qwen27BQ8_048 GB~30 GBYesqwen3.6:27b-q8_0
Qwen3 30BQwen30BQ4_K_M32 GB~22 GBYesqwen3:30b
Qwen3.5 35B-A3B InstructQwen35BQ4_K_M32 GB~20 GBYesqwen3.5:35b-a3b
Qwen3.6 35B-A3BQwen35BQ4_K_M32 GB~22 GBYesqwen3.6:35b-a3b
Qwen3.6 35B-A3B (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.6:35b-a3b-q8_0
Qwen3.5 35B-A3B Instruct (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.5:35b-a3b-q8_0
Qwen3-Next 80B-A3BQwen80BQ4_K_M72 GB~50.4 GBYesqwen3-next:80b
Qwen3-Next 80B-A3B (Q8)Qwen80BQ8_0128 GB~84.8 GBYesqwen3-next:80b-a3b-instruct-q8_0
Qwen3.5 122B-A10B InstructQwen122BQ4_K_M96 GB~72 GBYesqwen3.5:122b-a10b
Qwen3 235B A22BQwen235BQ4_K_M192 GB~130 GBYesqwen3:235b-a22b-q4_K_M
SmolLM2 360MSmolLM0.36BQ4_K_M1 GB~0.5 GBYessmollm2:360m
Grok 4.3xAIUndisclosedAPI0 GBCloud
GLM-4 PlusZhipuUndisclosedAPI0 GBCloud
GLM-5.2ZhipuUndisclosedAPI0 GBCloud
GLM-5Zhipu744BAPI0 GBOpen, no fit
GLM-5.1Zhipu744BAPI0 GBOpen, no fit

Estimated load = approximate memory at Q4_K_M; estimates, not measured. All local entries are GGUF builds pulled via Ollama, so they also run in llama.cpp and LM Studio. See the hardware stats for RAM-tier guidance.

Frequently asked questions

What does the ModelFit compatibility matrix show?

Every model ModelFit tracks (107 total, 75 local), with its parameter size, quantization, minimum RAM, and estimated memory load, so you can see at a glance which local AI models run on which Apple Silicon or NVIDIA hardware.

Can I download the dataset?

Yes. A machine-readable JSON export is free at /api/dataset/, licensed CC BY 4.0. The same data is also mirrored on GitHub and Hugging Face for offline use or bulk analysis.

Does "runs locally: false" mean a model is closed?

No. runsLocally tracks whether a registry-verified Ollama build fits a consumer RAM tier ModelFit maps (up to 256GB). Some open-weight models publish weights but exceed every tier. NVIDIA Nemotron 3 Ultra (550B, OpenMDW license) needs roughly 190 GB even at 2-bit quantization, so it is marked openWeights: true with runsLocally: false.

How is model fit calculated?

A model needs roughly 0.6 GB of memory per billion parameters at Q4_K_M quantization. ModelFit sizes recommendations to ~70% of a device's unified memory up to 32GB, scaling to ~85% at 128GB and above, leaving headroom for the OS, context, and KV-cache. On high-RAM Macs you can raise the GPU-wired ceiling further with iogpu.wired_limit_mb.