AI Model Families for Local Inference

Browse open-weight model families you can run locally with Ollama. Each family page shows all variants, RAM requirements, device compatibility, and performance expectations.

Catalog coverage

Quality-gated, not inflated

ModelFit tracks 141 models across 24 families, with 106 local rows. Every Ollama command is registry-verified before it ships, and low-quality 2-bit quants stay out of the ranking even when they technically exist.

141 tracked106 local7 quant classes105 registry-verified
Default policy
Q4_K_M is the default recommendation tier; Q6/Q8 variants appear when they add real quality headroom.
Freshness
Catalog snapshot updated 2026-09-03; new tags enter through the registry-check pipeline, not from scraped guesses.
Honest coverage
A smaller verified catalog beats a bigger list padded with broken tags and unusable quants.
Local-Runnable Models

AI Model Families for Local Use

Qwen logoAlibaba Cloud
Qwen

Qwen is Alibaba Cloud's open-weight model family with the widest range of sizes, from 0.5B to 235B parameters. Known for strong multilingual performance and coding ability.

34 models·0.5B–235B
Widest size range (0.5B to 235B)Strong multilingual and coding performance
Llama logoMeta
Llama

Llama is Meta's open-weight model family and the most popular choice for local AI. Known for strong general reasoning and a massive community ecosystem.

12 models·1B–405B
Most popular open-weight model familyStrong general reasoning and instruction following
DeepSeek logoDeepSeek AI
DeepSeek

DeepSeek specializes in reasoning and coding models. DeepSeek R1 introduced chain-of-thought reasoning that rivals proprietary models, while V3 is a massive MoE model.

6 models·7B–671B
Best-in-class reasoning with R1 modelsStrong coding performance
Mistral logoMistral AI
Mistral

Mistral AI's models are known for efficiency and strong performance relative to their size. Mistral 7B was a breakthrough that proved small models could compete with much larger ones.

9 models·7B–46.7B
Excellent performance-per-parameter ratioSliding window attention for efficiency
Gemma logoGoogle DeepMind
Gemma

Gemma is Google DeepMind's lightweight open model family. Known for excellent quality at small sizes and strong safety tuning.

19 models·1B–31B
Excellent quality at small sizes (1B-9B)Strong safety and instruction tuning
Phi logoMicrosoft
Phi

Phi is Microsoft's small-but-mighty model family, built on the idea that careful training data beats raw parameter count. Phi-4 Mini packs strong reasoning into just 3.8B parameters, while Phi-4 14B competes with much larger models on quality. Both run locally with Ollama and pair naturally with low-RAM Apple Silicon Macs.

6 models·3.8B–14B
Best quality-per-gigabyte at small sizesPhi-4 Mini 3.8B loads in just 3.2GB (7GB min RAM)
LFM2 logoLiquid AI
LFM2

LFM2 is Liquid AI's efficiency-focused model family, built on a hybrid architecture rather than a standard dense transformer. Its flagship, LFM2 24B-A2B, is a sparse mixture-of-experts model that activates only 2B of its 24B parameters per token. That design makes it fast on consumer hardware and well suited to agent workflows, tool calling, and privacy-sensitive local setups.

2 models·8.3B–24B
Hybrid MoE design: 24B total parameters, only 2B active per tokenLoads in about 14GB, fits any 16GB Mac
SmolLM logoHugging Face
SmolLM

SmolLM is Hugging Face's ultra-tiny model family for the most constrained devices. SmolLM2 360M loads in about 0.5GB and runs on anything with 1GB of RAM, from old Macs and iPhones to embedded boards. It is the smallest model in our database and the fastest by speed score.

1 models·0.36B–0.36B
Tiny 360M-parameter model from Hugging FaceLoads in about 0.5GB, needs just 1GB of RAM

Full Model Catalog

Filter every tracked model by RAM, family, runtime and context budget. Load figures and fit rules come from the same dataset that powers the wizard.

141 of 141 models
FamilyQuantLocalollama
Qwen2.5 1.5B InstructQwen1.5BQ4_K_M4 GB~1.5 GBYesqwen2.5:1.5b-instruct-q4_K_M
Qwen2.5 3B InstructQwen3BQ4_K_M4 GB~2.5 GBYesqwen2.5:3b-instruct-q4_K_M
Qwen2.5 7B InstructQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5:7b-instruct-q4_K_M
Qwen2.5 14B InstructQwen14BQ4_K_M16 GB~11 GBYesqwen2.5:14b-instruct-q4_K_M
Llama 3.2 3B InstructLlama3BQ4_K_M4 GB~2.5 GBYesllama3.2:3b-instruct-q4_K_M
Llama 3.1 8B InstructLlama8BQ4_K_M12 GB~6.5 GBYesllama3.1:8b-instruct-q4_K_M
Llama 3.1 8B Instruct (Q8)Llama8BQ8_016 GB~8 GBYesllama3.1:8b-instruct-q8_0
Llama 3.1 8B Instruct (Q5)Llama8BQ5_K_M12 GB~8 GBYesllama3.1:8b-instruct-q5_K_M
Llama 3.1 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.1:70b-instruct-q4_K_M
Mistral 7B InstructMistral7BQ4_K_M8 GB~5.5 GBYesmistral:7b-instruct-q4_K_M
Mistral 7B Instruct (Q5)Mistral7BQ5_K_M8 GB~4.8 GBYesmistral:7b-instruct-q5_K_M
Mistral 7B Instruct (Q8)Mistral7BQ8_016 GB~7.2 GBYesmistral:7b-instruct-q8_0
Mixtral 8x7B InstructMistral46.7BQ4_K_M48 GB~30 GBYesmixtral:8x7b
Mistral Nemo 12BMistral12BQ4_K_M16 GB~9.5 GBYesmistral-nemo:12b
Mistral Nemo 12B (Q8)Mistral12BQ8_024 GB~12.1 GBYesmistral-nemo:12b-instruct-2407-q8_0
Gemma 2 2B InstructGemma2BQ4_K_M4 GB~1.8 GBYesgemma2:2b-instruct-q4_K_M
Gemma 2 9B InstructGemma9BQ4_K_M12 GB~7 GBYesgemma2:9b-instruct-q4_K_M
Gemma 2 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma2:27b-instruct-q4_K_M
Phi-3 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi3:mini
Phi-3 Medium 14BPhi14BQ4_K_M16 GB~11 GBYesphi3:medium
Phi-4 14BPhi14BQ4_K_M24 GB~11.5 GBYesphi4:14b-q4_K_M
Phi-4 14B (Q8)Phi14BQ8_024 GB~14.5 GBYesphi4:14b-q8_0
Qwen2.5 Coder 7BQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5-coder:7b
Qwen2.5 Coder 14BQwen14BQ4_K_M16 GB~11 GBYesqwen2.5-coder:14b
Llama 3.2 1B InstructLlama1BQ4_K_M2 GB~1 GBYesllama3.2:1b-instruct-q4_K_M
Mistral Small 22BMistral22BQ4_K_M32 GB~17 GBYesmistral-small:22b
Kimi K2 InstructKimi1000BAPI0 GBOpen, no fit
Claude 3.5 SonnetClaudeUndisclosedAPI0 GBCloud
Claude 3.7 SonnetClaudeUndisclosedAPI0 GBCloud
Claude 3 OpusClaudeUndisclosedAPI0 GBCloud
Claude 4 OpusClaudeUndisclosedAPI0 GBCloud
Qwen2.5 0.5B InstructQwen0.5BQ4_K_M2 GB~0.8 GBYesqwen2.5:0.5b-instruct-q4_K_M
Gemma 3 1B InstructGemma1BQ4_K_M2 GB~1 GBYesgemma3:1b
Gemma 3 1B Instruct (Q8)Gemma1BQ8_04 GB~1 GBYesgemma3:1b-it-q8_0
Phi-4 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi4-mini:3.8b
Phi-4 Mini 3.8B (Q8)Phi3.8BQ8_08 GB~3.8 GBYesphi4-mini:3.8b-q8_0
SmolLM2 360MSmolLM0.36BQ4_K_M1 GB~0.5 GBYessmollm2:360m
Llama 3.3 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.3:70b-instruct-q4_K_M
Llama 3.1 405B InstructLlama405BQ4_K_M320 GB~243 GBYesllama3.1:405b-instruct-q4_K_M
DeepSeek-R1 671BDeepSeek671BQ4_K_M512 GB~380 GBYesdeepseek-r1:671b-q4_K_M
Qwen3 8BQwen8BQ4_K_M12 GB~6.5 GBYesqwen3:8b-q4_K_M
Qwen3 8B (Q8)Qwen8BQ8_016 GB~8.1 GBYesqwen3:8b-q8_0
Qwen3 14BQwen14BQ4_K_M16 GB~11 GBYesqwen3:14b-q4_K_M
Qwen3 30BQwen30BQ4_K_M32 GB~22 GBYesqwen3:30b
Qwen3 30B (Q8)Qwen30BQ8_048 GB~30.3 GBYesqwen3:30b-a3b-q8_0
Gemma 3 4B InstructGemma4BQ4_K_M6 GB~3.5 GBYesgemma3:4b
Gemma 3 4B Instruct (Q8)Gemma4BQ8_08 GB~3.9 GBYesgemma3:4b-it-q8_0
Gemma 3 12B InstructGemma12BQ4_K_M16 GB~9.5 GBYesgemma3:12b
Gemma 3 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma3:27b
DeepSeek-R1 Distill Qwen 7BDeepSeek7BQ4_K_M8 GB~5.5 GBYesdeepseek-r1:7b
DeepSeek-R1 Distill Qwen 7B (Q8)DeepSeek7BQ8_016 GB~7.5 GBYesdeepseek-r1:7b-qwen-distill-q8_0
DeepSeek-R1 Distill Qwen 14BDeepSeek14BQ4_K_M16 GB~11 GBYesdeepseek-r1:14b
DeepSeek-R1 Distill Qwen 14B (Q8)DeepSeek14BQ8_024 GB~14.6 GBYesdeepseek-r1:14b-qwen-distill-q8_0
DeepSeek-R1 Distill Llama 70BDeepSeek70BQ4_K_M64 GB~42 GBYesdeepseek-r1:70b
GPT-4oOpenAIUndisclosedAPI0 GBCloud
GPT-4o miniOpenAIUndisclosedAPI0 GBCloud
Claude 4 SonnetClaudeUndisclosedAPI0 GBCloud
Gemini 2.5 ProGoogleUndisclosedAPI0 GBCloud
Gemini 2.5 FlashGoogleUndisclosedAPI0 GBCloud
DeepSeek-V3DeepSeek671BAPI0 GBOpen, no fit
GLM-5Zhipu744BAPI0 GBOpen, no fit
GLM-4 PlusZhipuUndisclosedAPI0 GBCloud
Qwen3 235B A22BQwen235BQ4_K_M192 GB~130 GBYesqwen3:235b-a22b-q4_K_M
Mistral Small 3.1Mistral24BQ4_K_M24 GB~15 GBYesmistral-small3.1:24b
Mistral Small 3.1 (Q8)Mistral24BQ8_036 GB~23.3 GBYesmistral-small3.1:24b-instruct-2503-q8_0
DeepSeek-V3-0324DeepSeek671BAPI0 GBOpen, no fit
DeepSeek-R1DeepSeek671BAPI0 GBOpen, no fit
Qwen3.5 0.8B InstructQwen0.8BQ4_K_M2 GB~0.8 GBYesqwen3.5:0.8b
Qwen3.5 2B InstructQwen2BQ4_K_M4 GB~1.8 GBYesqwen3.5:2b
Qwen3.5 4B InstructQwen4BQ4_K_M6 GB~3.5 GBYesqwen3.5:4b
Qwen3.5 4B Instruct (Q8)Qwen4BQ8_08 GB~4.3 GBYesqwen3.5:4b-q8_0
Qwen3.5 9B InstructQwen9BQ4_K_M12 GB~7 GBYesqwen3.5:9b
Qwen3.5 35B-A3B InstructQwen35BQ4_K_M32 GB~20 GBYesqwen3.5:35b-a3b
Qwen3.5 27B InstructQwen27BQ4_K_M24 GB~16 GBYesqwen3.5:27b
Qwen3.5 27B Instruct (Q8)Qwen27BQ8_048 GB~27.1 GBYesqwen3.5:27b-q8_0
Qwen3.5 122B-A10B InstructQwen122BQ4_K_M96 GB~72 GBYesqwen3.5:122b-a10b
LFM2 24B-A2B InstructLFM224BQ4_K_M24 GB~14 GBYeslfm2:24b-a2b
LFM2.5 8B-A1BLFM28.3BQ4_K_M8 GB~5.5 GBYeslfm2.5:8b-a1b-q4_K_M
Granite 4.1 3B InstructGranite3BQ4_K_M4 GB~2 GBYesgranite4.1:3b
Granite 4.1 8B InstructGranite8BQ4_K_M8 GB~5.5 GBYesgranite4.1:8b
Qwen3.6 27BQwen27BQ4_K_M32 GB~18 GBYesqwen3.6:27b
Qwen3.8 27BQwen27BQ4_K_M24 GB~16.5 GBYesqwen3.8:27b
Qwen3.8 27B (Q8)Qwen27BQ8_048 GB~27.1 GBYesqwen3.8:27b-q8_0
Qwen3.6 35B-A3BQwen35BQ4_K_M32 GB~22 GBYesqwen3.6:35b-a3b
Gemma 4 31BGemma31BQ4_K_M32 GB~20 GBYesgemma4:31b
Gemma 4 31B (Q8)Gemma31BQ8_048 GB~30.9 GBYesgemma4:31b-it-q8_0
Gemma 4 26B-A4BGemma26BQ4_K_M24 GB~16 GBYesgemma4:26b
Gemma 4 E4BGemma4.5BQ4_K_M6 GB~4 GBYesgemma4:e4b
Gemma 4 E4B (Q8)Gemma4.5BQ8_016 GB~7.5 GBYesgemma4:e4b-it-q8_0
Gemma 4 E2BGemma2.3BQ4_K_M4 GB~2.3 GBYesgemma4:e2b
Gemma 4 E2B (Q8)Gemma2.3BQ8_08 GB~4.6 GBYesgemma4:e2b-it-q8_0
Llama 4 ScoutLlama109BQ4_K_M96 GB~67 GBYesllama4:scout
Llama 4 MaverickLlama400BQ4_K_M320 GB~245 GBYesllama4:maverick
Mistral Medium 3.5Mistral128BAPI0 GBOpen (llama.cpp)
DeepSeek V4 Flash 0731DeepSeek284BAPI0 GBOpen (llama.cpp)
DeepSeek V4 ProDeepSeek1600BAPI0 GBOpen, no fit
Kimi K2.6Kimi1000BAPI0 GBOpen, no fit
GLM-5.1Zhipu744BAPI0 GBOpen, no fit
Claude Opus 4.7ClaudeUndisclosedAPI0 GBCloud
Claude Opus 4.8ClaudeUndisclosedAPI0 GBCloud
GPT-5.5OpenAIUndisclosedAPI0 GBCloud
Gemini 3.1 ProGoogleUndisclosedAPI0 GBCloud
Grok 4.3xAIUndisclosedAPI0 GBCloud
Gemma 4 12BGemma12BQ4_K_M12 GB~8 GBYesgemma4:12b
Claude Fable 5ClaudeUndisclosedAPI0 GBCloud
GLM-5.2Zhipu753BAPI0 GBOpen, no fit
Kimi K2.7-CodeKimi1000BAPI0 GBOpen, no fit
MiniMax M3MiniMax428BAPI0 GBOpen, no fit
Qwen3.7-PlusQwenUndisclosedAPI0 GBCloud
NVIDIA Nemotron 3 UltraNemotron550BAPI0 GBOpen, no fit
Xiaomi MiMo-V2-FlashMiMo309BAPI0 GBOpen (llama.cpp)
NVIDIA Nemotron Cascade 2 30B-A3BNemotron30BQ6_K36 GB~24 GBYesnemotron-cascade-2:30b
Poolside Laguna XS.2Laguna33BQ4_K_M36 GB~23 GBYeslaguna-xs.2:q4_K_M
Cohere North Mini CodeNorth30BQ4_K_M32 GB~19 GBYesnorth-mini-code-1.0:q4_K_M
GPT-OSS 120BGPT-OSS117BMXFP496 GB~65.4 GBYesgpt-oss:120b
GPT-OSS 20BGPT-OSS21BMXFP424 GB~13.8 GBYesgpt-oss:20b
Qwen3-Next 80B-A3BQwen80BQ4_K_M72 GB~50.4 GBYesqwen3-next:80b
Qwen3.6 35B-A3B (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.6:35b-a3b-q8_0
Qwen3.5 35B-A3B Instruct (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.5:35b-a3b-q8_0
Qwen3.5 9B Instruct (Q8)Qwen9BQ8_016 GB~10.7 GBYesqwen3.5:9b-q8_0
Qwen3.6 27B (Q8)Qwen27BQ8_048 GB~30 GBYesqwen3.6:27b-q8_0
Gemma 4 12B (Q8)Gemma12BQ8_024 GB~12.8 GBYesgemma4:12b-it-q8_0
Gemma 4 26B-A4B (Q8)Gemma26BQ8_048 GB~28.1 GBYesgemma4:26b-a4b-it-q8_0
Qwen3-Next 80B-A3B (Q8)Qwen80BQ8_0128 GB~84.8 GBYesqwen3-next:80b-a3b-instruct-q8_0
Qwen3 14B (Q8)Qwen14BQ8_024 GB~15.9 GBYesqwen3:14b-q8_0
Llama 3.3 70B Instruct (Q8)Llama70BQ8_096 GB~75 GBYesllama3.3:70b-instruct-q8_0
Llama 3.3 70B Instruct (Q6)Llama70BQ6_K96 GB~57.9 GBYesllama3.3:70b-instruct-q6_K
Kimi K3Kimi2800BAPI0 GBOpen, no fit
Laguna XS 2.1Laguna33BQ4_K_M32 GB~20.3 GBYeslaguna-xs-2.1:q4_K_M
Laguna S 2.1Laguna118BQ4_K_M128 GB~96 GBYeslaguna-s-2.1:q4_K_M
Ornith 1.0 9BOrnith9BQ4_K_M8 GB~5.6 GBYesornith:9b
Ornith 1.0 35BOrnith35BQ4_K_M32 GB~21.2 GBYesornith:35b
Qwen3.8-Flash-NextQwen125BIQ1_S192 GB~123 GBYes
Granite 4.2 3BGranite3.7BQ4_K_M4 GB~2.1 GBYesgranite4.2:3b
Granite 4.2 8BGranite8.8BQ4_K_M8 GB~5 GBYesgranite4.2:8b
Granite 4.2 30BGranite29.3BQ4_K_M24 GB~16.5 GBYesgranite4.2:30b
Nemotron 3.5 Lightning 30B-A3BNemotron30BQ4_K_M36 GB~23.7 GBYesnemotron-3.5-lightning:30b
Muse Glimmer 30BMuse29.8BQ4_K_M24 GB~16.9 GBYesmuse-glimmer:30b
MiniCPM-V 4.5 8BMiniCPM8.7BQ4_K_M8 GB~5.7 GBYesminicpm-v4.5:8b
GLM-5.3Zhipu753BAPI0 GBOpen, no fit
GLM-5.3 FlashZhipu321BAPI0 GBOpen (llama.cpp)

How Much RAM Do You Need?

Bar chart: maximum local LLM size by memory tier. 8 GB runs up to 9B, 12 GB runs up to 12B, 16 GB runs up to 14B, 24 GB runs up to 27B, 32 GB runs up to 35B, 36 GB runs up to 35B, 48 GB runs up to 35B, 64 GB runs up to 70B, 72 GB runs up to 70B, 96 GB runs up to 70B, 128 GB runs up to 70B, 192 GB runs up to 70B, 256 GB runs up to 70B, 512 GB runs up to 405B. Data from ModelFit's own catalog.

Largest dense Q4 model that fits each memory tier, from ModelFit’s own catalog. Full breakdown on the stats page.

Explore by Hardware