AI Model Families for Local Inference

Browse open-weight model families you can run locally with Ollama. Each family page shows all variants, RAM requirements, device compatibility, and performance expectations.

Catalog coverage

Quality-gated, not inflated

ModelFit tracks 152 models across 26 families, with 113 local rows. Every Ollama command is registry-verified before it ships, and low-quality 2-bit quants stay out of the ranking even when they technically exist.

152 tracked113 local7 quant classes110 registry-verified
Default policy
Q4_K_M is the default recommendation tier; Q6/Q8 variants appear when they add real quality headroom.
Freshness
Catalog snapshot updated 2026-09-29; new tags enter through the registry-check pipeline, not from scraped guesses.
Honest coverage
A smaller verified catalog beats a bigger list padded with broken tags and unusable quants.
Local-Runnable Models

AI Model Families for Local Use

Qwen logoAlibaba Cloud
Qwen

Qwen is Alibaba Cloud's open-weight model family spanning 0.5B to 2.4T parameters. Qwen3.8 27B is the local flagship: a dense, Apache 2.0 vision-language model, while the Qwen3.8 2.4T A95B MoE tops the cloud line.

34 models·0.5B–235B
Qwen3.8 27B local flagship (Apache 2.0, vision)Cloud line tops out at the 2.4T A95B MoE
Llama logoMeta
Llama

Llama is Meta's open-weight model family and the most popular choice for local AI. The Llama 4 line brings multimodal MoE models for big-memory machines, while the 1B-70B Llama 3.x sizes remain the practical laptop and desktop picks.

12 models·1B–405B
Most popular open-weight model familyStrong general reasoning and instruction following
DeepSeek logoDeepSeek AI
DeepSeek

DeepSeek now ships two distinct lines. The V4 generation is cloud-first: V4.1 Flash is an MIT-licensed multimodal image + text MoE, and V4 Flash 0731 and V4 Pro are API models. On consumer hardware DeepSeek means the R1 reasoning distills, from 7B up to the full 671B, all registry-verified in Ollama.

6 models·7B–671B
V4.1 Flash: multimodal 552B MoE, API onlyR1 distills 7B-70B run locally in Ollama
Mistral logoMistral AI
Mistral

Mistral AI's models are known for efficiency and strong performance relative to their size. Mistral 7B was a breakthrough that proved small models could compete with much larger ones.

11 models·7B–123B
Excellent performance-per-parameter ratioSliding window attention for efficiency
Gemma logoGoogle DeepMind
Gemma

Gemma is Google DeepMind's efficient open model family, tuned for strong quality at small sizes. The Gemma 4 line spans an on-device 4.5B model up to a dense 31B and a 26B MoE, all Apache 2.0.

19 models·1B–31B
Gemma 4 E4B: 4.5B effective, runs on 8GB MacsGemma 4 26B-A4B MoE: 25.2B total, 3.8B active
Phi logoMicrosoft
Phi

Phi is Microsoft's small-but-mighty model family, built on the idea that careful training data beats raw parameter count. Phi-4 Mini packs strong reasoning into just 3.8B parameters, while Phi-4 14B competes with much larger models on quality. Both run locally with Ollama and pair naturally with low-RAM Apple Silicon Macs.

6 models·3.8B–14B
Best quality-per-gigabyte at small sizesPhi-4 Mini 3.8B loads in just 3.2GB (7GB min RAM)
LFM2 logoLiquid AI
LFM2

LFM2 is Liquid AI's efficiency-focused model family, built on a hybrid architecture rather than a standard dense transformer. Its flagship, LFM2 24B-A2B, is a sparse mixture-of-experts model that activates only 2B of its 24B parameters per token. That design makes it fast on consumer hardware and well suited to agent workflows, tool calling, and privacy-sensitive local setups.

2 models·8.3B–24B
Hybrid MoE design: 24B total parameters, only 2B active per tokenLoads in about 14GB, fits any 16GB Mac
SmolLM logoHugging Face
SmolLM

SmolLM is Hugging Face's ultra-tiny model family for the most constrained devices. SmolLM2 360M loads in about 0.5GB and runs on anything with 1GB of RAM, from old Macs and iPhones to embedded boards. It is the smallest model in our database and the fastest by speed score.

1 models·0.36B–0.36B
Tiny 360M-parameter model from Hugging FaceLoads in about 0.5GB, needs just 1GB of RAM
Beyond Consumer Hardware

API-only & open-weights flagships

Frontier models you don't run on a laptop — hosted APIs, often with open weights. Each page says what it does and lists the local alternatives that actually fit your machine.

Full Model Catalog

Filter every tracked model by RAM, family, runtime and context budget. Load figures and fit rules come from the same dataset that powers the wizard.

152 of 152 models
FamilyQuantLocalollama
Qwen2.5 1.5B InstructQwen1.5BQ4_K_M4 GB~1.5 GBYesqwen2.5:1.5b-instruct-q4_K_M
Qwen2.5 3B InstructQwen3BQ4_K_M4 GB~2.5 GBYesqwen2.5:3b-instruct-q4_K_M
Qwen2.5 7B InstructQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5:7b-instruct-q4_K_M
Qwen2.5 14B InstructQwen14BQ4_K_M16 GB~11 GBYesqwen2.5:14b-instruct-q4_K_M
Llama 3.2 3B InstructLlama3BQ4_K_M4 GB~2.5 GBYesllama3.2:3b-instruct-q4_K_M
Llama 3.1 8B InstructLlama8BQ4_K_M12 GB~6.5 GBYesllama3.1:8b-instruct-q4_K_M
Llama 3.1 8B Instruct (Q8)Llama8BQ8_016 GB~8 GBYesllama3.1:8b-instruct-q8_0
Llama 3.1 8B Instruct (Q5)Llama8BQ5_K_M12 GB~8 GBYesllama3.1:8b-instruct-q5_K_M
Llama 3.1 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.1:70b-instruct-q4_K_M
Mistral 7B InstructMistral7BQ4_K_M8 GB~5.5 GBYesmistral:7b-instruct-q4_K_M
Mistral 7B Instruct (Q5)Mistral7BQ5_K_M8 GB~4.8 GBYesmistral:7b-instruct-q5_K_M
Mistral 7B Instruct (Q8)Mistral7BQ8_016 GB~7.2 GBYesmistral:7b-instruct-q8_0
Mixtral 8x7B InstructMistral46.7BQ4_K_M48 GB~30 GBYesmixtral:8x7b
Mistral Nemo 12BMistral12BQ4_K_M16 GB~9.5 GBYesmistral-nemo:12b
Mistral Nemo 12B (Q8)Mistral12BQ8_024 GB~12.1 GBYesmistral-nemo:12b-instruct-2407-q8_0
Gemma 2 2B InstructGemma2BQ4_K_M4 GB~1.8 GBYesgemma2:2b-instruct-q4_K_M
Gemma 2 9B InstructGemma9BQ4_K_M12 GB~7 GBYesgemma2:9b-instruct-q4_K_M
Gemma 2 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma2:27b-instruct-q4_K_M
Phi-3 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi3:mini
Phi-3 Medium 14BPhi14BQ4_K_M16 GB~11 GBYesphi3:medium
Phi-4 14BPhi14BQ4_K_M24 GB~11.5 GBYesphi4:14b-q4_K_M
Phi-4 14B (Q8)Phi14BQ8_024 GB~14.5 GBYesphi4:14b-q8_0
Qwen2.5 Coder 7BQwen7BQ4_K_M8 GB~5.5 GBYesqwen2.5-coder:7b
Qwen2.5 Coder 14BQwen14BQ4_K_M16 GB~11 GBYesqwen2.5-coder:14b
Llama 3.2 1B InstructLlama1BQ4_K_M2 GB~1 GBYesllama3.2:1b-instruct-q4_K_M
Mistral Small 22BMistral22BQ4_K_M32 GB~17 GBYesmistral-small:22b
Kimi K2 InstructKimi1000BAPI0 GB—Open, no fit—
Claude 3.5 SonnetClaudeUndisclosedAPI0 GB—Cloud—
Claude 3.7 SonnetClaudeUndisclosedAPI0 GB—Cloud—
Claude 3 OpusClaudeUndisclosedAPI0 GB—Cloud—
Claude 4 OpusClaudeUndisclosedAPI0 GB—Cloud—
Qwen2.5 0.5B InstructQwen0.5BQ4_K_M2 GB~0.8 GBYesqwen2.5:0.5b-instruct-q4_K_M
Gemma 3 1B InstructGemma1BQ4_K_M2 GB~1 GBYesgemma3:1b
Gemma 3 1B Instruct (Q8)Gemma1BQ8_04 GB~1 GBYesgemma3:1b-it-q8_0
Phi-4 Mini 3.8BPhi3.8BQ4_K_M6 GB~3.2 GBYesphi4-mini:3.8b
Phi-4 Mini 3.8B (Q8)Phi3.8BQ8_08 GB~3.8 GBYesphi4-mini:3.8b-q8_0
SmolLM2 360MSmolLM0.36BQ4_K_M1 GB~0.5 GBYessmollm2:360m
Llama 3.3 70B InstructLlama70BQ4_K_M64 GB~42 GBYesllama3.3:70b-instruct-q4_K_M
Llama 3.1 405B InstructLlama405BQ4_K_M320 GB~243 GBYesllama3.1:405b-instruct-q4_K_M
DeepSeek-R1 671BDeepSeek671BQ4_K_M512 GB~380 GBYesdeepseek-r1:671b-q4_K_M
Qwen3 8BQwen8BQ4_K_M12 GB~6.5 GBYesqwen3:8b-q4_K_M
Qwen3 8B (Q8)Qwen8BQ8_016 GB~8.1 GBYesqwen3:8b-q8_0
Qwen3 14BQwen14BQ4_K_M16 GB~11 GBYesqwen3:14b-q4_K_M
Qwen3 30BQwen30BQ4_K_M32 GB~22 GBYesqwen3:30b
Qwen3 30B (Q8)Qwen30BQ8_048 GB~30.3 GBYesqwen3:30b-a3b-q8_0
Gemma 3 4B InstructGemma4BQ4_K_M6 GB~3.5 GBYesgemma3:4b
Gemma 3 4B Instruct (Q8)Gemma4BQ8_08 GB~3.9 GBYesgemma3:4b-it-q8_0
Gemma 3 12B InstructGemma12BQ4_K_M16 GB~9.5 GBYesgemma3:12b
Gemma 3 27B InstructGemma27BQ4_K_M32 GB~21 GBYesgemma3:27b
DeepSeek-R1 Distill Qwen 7BDeepSeek7BQ4_K_M8 GB~5.5 GBYesdeepseek-r1:7b
DeepSeek-R1 Distill Qwen 7B (Q8)DeepSeek7BQ8_016 GB~7.5 GBYesdeepseek-r1:7b-qwen-distill-q8_0
DeepSeek-R1 Distill Qwen 14BDeepSeek14BQ4_K_M16 GB~11 GBYesdeepseek-r1:14b
DeepSeek-R1 Distill Qwen 14B (Q8)DeepSeek14BQ8_024 GB~14.6 GBYesdeepseek-r1:14b-qwen-distill-q8_0
DeepSeek-R1 Distill Llama 70BDeepSeek70BQ4_K_M64 GB~42 GBYesdeepseek-r1:70b
GPT-4oOpenAIUndisclosedAPI0 GB—Cloud—
GPT-4o miniOpenAIUndisclosedAPI0 GB—Cloud—
Claude 4 SonnetClaudeUndisclosedAPI0 GB—Cloud—
Gemini 2.5 ProGoogleUndisclosedAPI0 GB—Cloud—
Gemini 2.5 FlashGoogleUndisclosedAPI0 GB—Cloud—
DeepSeek-V3DeepSeek671BAPI0 GB—Open, no fit—
GLM-5Zhipu744BAPI0 GB—Open, no fit—
GLM-4 PlusZhipuUndisclosedAPI0 GB—Cloud—
Qwen3 235B A22BQwen235BQ4_K_M192 GB~130 GBYesqwen3:235b-a22b-q4_K_M
Mistral Small 3.1Mistral24BQ4_K_M24 GB~15 GBYesmistral-small3.1:24b
Mistral Small 3.1 (Q8)Mistral24BQ8_036 GB~23.3 GBYesmistral-small3.1:24b-instruct-2503-q8_0
DeepSeek-V3-0324DeepSeek671BAPI0 GB—Open, no fit—
DeepSeek-R1DeepSeek671BAPI0 GB—Open, no fit—
Qwen3.5 0.8B InstructQwen0.8BQ4_K_M2 GB~0.8 GBYesqwen3.5:0.8b
Qwen3.5 2B InstructQwen2BQ4_K_M4 GB~1.8 GBYesqwen3.5:2b
Qwen3.5 4B InstructQwen4BQ4_K_M6 GB~3.5 GBYesqwen3.5:4b
Qwen3.5 4B Instruct (Q8)Qwen4BQ8_08 GB~4.3 GBYesqwen3.5:4b-q8_0
Qwen3.5 9B InstructQwen9BQ4_K_M12 GB~7 GBYesqwen3.5:9b
Qwen3.5 35B-A3B InstructQwen35BQ4_K_M32 GB~20 GBYesqwen3.5:35b-a3b
Qwen3.5 27B InstructQwen27BQ4_K_M24 GB~16 GBYesqwen3.5:27b
Qwen3.5 27B Instruct (Q8)Qwen27BQ8_048 GB~27.1 GBYesqwen3.5:27b-q8_0
Qwen3.5 122B-A10B InstructQwen122BQ4_K_M96 GB~72 GBYesqwen3.5:122b-a10b
LFM2 24B-A2B InstructLFM224BQ4_K_M24 GB~14 GBYeslfm2:24b-a2b
LFM2.5 8B-A1BLFM28.3BQ4_K_M8 GB~5.5 GBYeslfm2.5:8b-a1b-q4_K_M
Granite 4.1 3B InstructGranite3BQ4_K_M4 GB~2 GBYesgranite4.1:3b
Granite 4.1 8B InstructGranite8BQ4_K_M8 GB~5.5 GBYesgranite4.1:8b
Qwen3.6 27BQwen27BQ4_K_M32 GB~18 GBYesqwen3.6:27b
Qwen3.8 27BQwen27BQ4_K_M24 GB~16.5 GBYesqwen3.8:27b
Qwen3.8 27B (Q8)Qwen27BQ8_048 GB~27.1 GBYesqwen3.8:27b-q8_0
Qwen3.6 35B-A3BQwen35BQ4_K_M32 GB~22 GBYesqwen3.6:35b-a3b
Gemma 4 31BGemma31BQ4_K_M32 GB~20 GBYesgemma4:31b
Gemma 4 31B (Q8)Gemma31BQ8_048 GB~30.9 GBYesgemma4:31b-it-q8_0
Gemma 4 26B-A4BGemma26BQ4_K_M24 GB~16 GBYesgemma4:26b
Gemma 4 E4BGemma4.5BQ4_K_M6 GB~4 GBYesgemma4:e4b
Gemma 4 E4B (Q8)Gemma4.5BQ8_016 GB~7.5 GBYesgemma4:e4b-it-q8_0
Gemma 4 E2BGemma2.3BQ4_K_M4 GB~2.3 GBYesgemma4:e2b
Gemma 4 E2B (Q8)Gemma2.3BQ8_08 GB~4.6 GBYesgemma4:e2b-it-q8_0
Llama 4 ScoutLlama109BQ4_K_M96 GB~67 GBYesllama4:scout
Llama 4 MaverickLlama400BQ4_K_M320 GB~245 GBYesllama4:maverick
Mistral Medium 3.5Mistral128BAPI0 GB—Open (llama.cpp)—
DeepSeek V4 Flash 0731DeepSeek284BAPI0 GB—Open (llama.cpp)—
DeepSeek V4 ProDeepSeek1600BAPI0 GB—Open, no fit—
DeepSeek V4.1 FlashDeepSeek552BAPI0 GB—Open, no fit—
Kimi K2.6Kimi1000BAPI0 GB—Open, no fit—
GLM-5.1Zhipu744BAPI0 GB—Open, no fit—
Claude Opus 4.7ClaudeUndisclosedAPI0 GB—Cloud—
Claude Opus 4.8ClaudeUndisclosedAPI0 GB—Cloud—
GPT-5.5OpenAIUndisclosedAPI0 GB—Cloud—
Gemini 3.1 ProGoogleUndisclosedAPI0 GB—Cloud—
Grok 4.3xAIUndisclosedAPI0 GB—Cloud—
Gemma 4 12BGemma12BQ4_K_M12 GB~8 GBYesgemma4:12b
Claude Fable 5ClaudeUndisclosedAPI0 GB—Cloud—
GLM-5.2Zhipu753BAPI0 GB—Open, no fit—
Kimi K2.7-CodeKimi1000BAPI0 GB—Open, no fit—
MiniMax M3MiniMax428BAPI0 GB—Open, no fit—
Qwen3.7-PlusQwenUndisclosedAPI0 GB—Cloud—
NVIDIA Nemotron 3 UltraNemotron550BAPI0 GB—Open, no fit—
Xiaomi MiMo-V2-FlashMiMo309BAPI0 GB—Open (llama.cpp)—
NVIDIA Nemotron Cascade 2 30B-A3BNemotron30BQ6_K36 GB~24 GBYesnemotron-cascade-2:30b
Poolside Laguna XS.2Laguna33BQ4_K_M36 GB~23 GBYeslaguna-xs.2:q4_K_M
Cohere North Mini CodeNorth30BQ4_K_M32 GB~19 GBYesnorth-mini-code-1.0:q4_K_M
GPT-OSS 120BGPT-OSS117BMXFP496 GB~65.4 GBYesgpt-oss:120b
GPT-OSS 20BGPT-OSS21BMXFP424 GB~13.8 GBYesgpt-oss:20b
Qwen3-Next 80B-A3BQwen80BQ4_K_M72 GB~50.4 GBYesqwen3-next:80b
Qwen3.6 35B-A3B (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.6:35b-a3b-q8_0
Qwen3.5 35B-A3B Instruct (Q8)Qwen35BQ8_064 GB~38.7 GBYesqwen3.5:35b-a3b-q8_0
Qwen3.5 9B Instruct (Q8)Qwen9BQ8_016 GB~10.7 GBYesqwen3.5:9b-q8_0
Qwen3.6 27B (Q8)Qwen27BQ8_048 GB~30 GBYesqwen3.6:27b-q8_0
Gemma 4 12B (Q8)Gemma12BQ8_024 GB~12.8 GBYesgemma4:12b-it-q8_0
Gemma 4 26B-A4B (Q8)Gemma26BQ8_048 GB~28.1 GBYesgemma4:26b-a4b-it-q8_0
Qwen3-Next 80B-A3B (Q8)Qwen80BQ8_0128 GB~84.8 GBYesqwen3-next:80b-a3b-instruct-q8_0
Qwen3 14B (Q8)Qwen14BQ8_024 GB~15.9 GBYesqwen3:14b-q8_0
Llama 3.3 70B Instruct (Q8)Llama70BQ8_096 GB~75 GBYesllama3.3:70b-instruct-q8_0
Llama 3.3 70B Instruct (Q6)Llama70BQ6_K96 GB~57.9 GBYesllama3.3:70b-instruct-q6_K
Kimi K3Kimi2800BAPI0 GB—Open, no fit—
Laguna XS 2.1Laguna33BQ4_K_M32 GB~20.3 GBYeslaguna-xs-2.1:q4_K_M
Laguna S 2.1Laguna118BQ4_K_M128 GB~96 GBYeslaguna-s-2.1:q4_K_M
Ornith 1.0 9BOrnith9BQ4_K_M8 GB~5.6 GBYesornith:9b
Ornith 1.0 35BOrnith35BQ4_K_M32 GB~21.2 GBYesornith:35b
Qwen3.8-Flash-NextQwen125BIQ1_S192 GB~123 GBYes—
Qwen3.8 2.4T A95BQwen2400BAPI0 GB—Open, no fit—
Granite 4.2 3BGranite3.7BQ4_K_M4 GB~2.1 GBYesgranite4.2:3b
Granite 4.2 8BGranite8.8BQ4_K_M8 GB~5 GBYesgranite4.2:8b
Granite 4.2 30BGranite29.3BQ4_K_M24 GB~16.5 GBYesgranite4.2:30b
Nemotron 3.5 Lightning 30B-A3BNemotron30BQ4_K_M36 GB~23.7 GBYesnemotron-3.5-lightning:30b
Muse Glimmer 30BMuse29.8BQ4_K_M32 GB~16.9 GBYesmuse-glimmer:30b
MiniCPM-V 4.5 8BMiniCPM8.7BQ4_K_M12 GB~5.7 GBYesminicpm-v4.5:8b
GLM-5.3Zhipu753BAPI0 GB—Open, no fit—
GLM-5.3 FlashZhipu321BAPI0 GB—Open (llama.cpp)—
GLM-4.7-Flash 30B-A3BZhipu30BQ4_K_M32 GB~19 GBYesglm-4.7-flash
Nemotron 3 Super 120B-A12BNemotron120BQ4_K_M128 GB~86.8 GBYesnemotron-3-super:120b
Olmo 3.1 32B InstructOlmo32BQ4_K_M32 GB~19.5 GBYesolmo-3.1:32b
MiMo-V2.6-Distill-Qwen-9BMiMo9BQ4_K_M16 GB~5.8 GBYes—
Ling-3.0-flashLing124BQ4_K_M128 GB~77.8 GBYes—
Xiaomi MiMo-V2.6-FlashMiMo309BAPI0 GB—Open (llama.cpp)—
Xiaomi MiMo-V2.6-ProMiMo1020BAPI0 GB—Open, no fit—
Devstral Small 2 24BMistral24BQ4_K_M24 GB~15.2 GBYesdevstral-small-2:24b
Devstral 2 123BMistral123BQ4_K_M96 GB~74.9 GBYesdevstral-2:123b

How Much RAM Do You Need?

Bar chart: maximum local LLM size by memory tier. 8 GB runs up to 9B, 12 GB runs up to 12B, 16 GB runs up to 14B, 24 GB runs up to 27B, 32 GB runs up to 35B, 36 GB runs up to 35B, 48 GB runs up to 35B, 64 GB runs up to 70B, 72 GB runs up to 70B, 96 GB runs up to 70B, 128 GB runs up to 70B, 192 GB runs up to 70B, 256 GB runs up to 70B, 512 GB runs up to 405B. Data from ModelFit's own catalog.

Largest dense Q4 model that fits each memory tier, from ModelFit’s own catalog. Full breakdown on the stats page.

Explore by Hardware