# ModelFit > Find the best local AI models for Apple Silicon Macs, iPhones, and NVIDIA GPUs. Interactive wizard, per-device rankings, and a registry-verified model catalog. ## Main Pages - [Home](https://modelfit.io/): Interactive wizard for local AI model recommendations by hardware - [Devices](https://modelfit.io/devices/): All supported Apple Silicon devices with ranked model recommendations - [Models](https://modelfit.io/models/): Browse model families (Qwen, Llama, Gemma, DeepSeek, and more) - [Benchmark](https://modelfit.io/benchmark/): Local LLMs vs cloud flagships with vendor-sourced SWE-Bench scores - [Hardware Stats](https://modelfit.io/stats/): RAM/VRAM per model tier, RAM-tier to max model size - [Compatibility Dataset](https://modelfit.io/data/): Open CC BY 4.0 dataset of models × hardware (JSON export at https://modelfit.io/api/dataset/) - [Blog](https://modelfit.io/blog/): Local LLM news, buyer guides, and Apple Silicon AI performance (RSS: https://modelfit.io/feed.xml) - [About & Disclosure](https://modelfit.io/about/): Methodology, estimate policy, and affiliate disclosure ## Key Facts > Free to cite with attribution to ModelFit (modelfit.io). Full list: https://modelfit.io/stats/ - ModelFit tracks 141 AI models across 24 families; 106 run locally via Ollama on Apple Silicon or NVIDIA GPUs (ModelFit, 2026). - At Q4_K_M quantization, a local LLM needs roughly 0.6 GB of memory per billion parameters for the weights, rising to about 0.8 GB per billion for models under 8B once runtime overhead is counted (ModelFit, 2026). - ModelFit sizes recommendations to ~70% of a device’s unified memory up to 32GB, scaling to ~85% at 128GB and above, leaving headroom for the OS, context, and KV-cache (ModelFit, 2026). - An 8GB device comfortably runs dense local models up to ~9B parameters at Q4; 32 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 12GB device comfortably runs dense local models up to ~12B parameters at Q4; 44 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 16GB device comfortably runs dense local models up to ~14B parameters at Q4; 52 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 24GB device comfortably runs dense local models up to ~29.3B parameters at Q4; 65 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 32GB device comfortably runs dense local models up to ~35B parameters at Q4; 77 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 36GB device comfortably runs dense local models up to ~35B parameters at Q4; 81 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 48GB device comfortably runs dense local models up to ~35B parameters at Q4, or up to ~46.7B total as a mixture-of-experts model; 88 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 64GB device comfortably runs dense local models up to ~70B parameters at Q4; 93 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 72GB device comfortably runs dense local models up to ~70B parameters at Q4, or up to ~80B total as a mixture-of-experts model; 94 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 96GB device comfortably runs dense local models up to ~70B parameters at Q4, or up to ~122B total as a mixture-of-experts model; 99 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 128GB device comfortably runs dense local models up to ~70B parameters at Q4, or up to ~122B total as a mixture-of-experts model; 101 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 192GB device comfortably runs dense local models up to ~70B parameters at Q4, or up to ~235B total as a mixture-of-experts model; 103 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 256GB device comfortably runs dense local models up to ~70B parameters at Q4, or up to ~235B total as a mixture-of-experts model; 103 of ModelFit’s 106 local models fit (ModelFit, 2026). - A 512GB device comfortably runs dense local models up to ~405B parameters at Q4, or up to ~671B total as a mixture-of-experts model; 106 of ModelFit’s 106 local models fit (ModelFit, 2026). ## Guides - [Install Ollama on Mac](https://modelfit.io/guides/ollama-setup/): Step-by-step setup and first model - [Best LLM for MacBook](https://modelfit.io/guides/best-llm-for-macbook/): Picks by chip generation and RAM tier - [How Much RAM for a Local LLM](https://modelfit.io/guides/how-much-ram-for-local-llm/): Model-size-to-memory matrix. ~0.6 GB per billion params at Q4, by RAM tier - [Best LLM for iPhone](https://modelfit.io/guides/best-llm-for-iphone/): On-device apps and models that fit 6-12GB - [Run AI Offline](https://modelfit.io/guides/run-ai-offline/): Fully offline local AI and air-gapped setups - [Best Local LLM for Coding](https://modelfit.io/guides/best-local-llm-for-coding/): Coding-model tier list by RAM with editor setup - [Local Image Generation on Mac](https://modelfit.io/guides/local-image-generation-mac/): FLUX/SDXL/SD 3.5 by Mac RAM tier with verified file sizes - [Does Ollama Use GPU or CPU](https://modelfit.io/guides/does-ollama-use-gpu/): GPU auto-detection, VRAM offload splits, and how to check with ollama ps - [Is Local AI Actually Private](https://modelfit.io/guides/is-local-ai-private/): What stays on-device, what touches the network - [Can an 8GB Mac Run an LLM](https://modelfit.io/guides/can-8gb-mac-run-llm/): Engine-derived picks for an 8GB budget with KV-cache costs ## Device Pages - [MacBook Air](https://modelfit.io/macbook-air/): Apple M5, 16GB default - [MacBook Pro](https://modelfit.io/macbook-pro/): Apple M5 Pro, 48GB default - [Mac Mini](https://modelfit.io/mac-mini/): Apple M6, 16GB default - [Mac Studio](https://modelfit.io/mac-studio/): Apple M5 Max, 64GB default - [iPhone 15](https://modelfit.io/iphone-15/): Apple A16, 6GB default - [iPhone 15 Pro](https://modelfit.io/iphone-15-pro/): Apple A17 Pro, 8GB default - [iPhone 15 Pro Max](https://modelfit.io/iphone-15-pro-max/): Apple A17 Pro, 8GB default - [iPhone 16](https://modelfit.io/iphone-16/): Apple A18, 8GB default - [iPhone 16 Pro](https://modelfit.io/iphone-16-pro/): Apple A18 Pro, 8GB default - [iPhone 16 Pro Max](https://modelfit.io/iphone-16-pro-max/): Apple A18 Pro, 8GB default - [iPhone 16e](https://modelfit.io/iphone-16e/): Apple A18, 8GB default - [iPhone 17](https://modelfit.io/iphone-17/): Apple A19, 8GB default - [iPhone 17 Air](https://modelfit.io/iphone-17-air/): Apple A19, 12GB default - [iPhone 17 Pro](https://modelfit.io/iphone-17-pro/): Apple A19 Pro, 12GB default - [iPhone 17 Pro Max](https://modelfit.io/iphone-17-pro-max/): Apple A19 Pro, 12GB default ## GPU Pages - [GPU Hub](https://modelfit.io/gpu/): NVIDIA GPUs ranked for local LLMs - [RTX 4060](https://modelfit.io/gpu/rtx-4060/) - [RTX 3060](https://modelfit.io/gpu/rtx-3060/) - [RTX 4060 Ti](https://modelfit.io/gpu/rtx-4060-ti/) - [RTX 5060 Ti](https://modelfit.io/gpu/rtx-5060-ti/) - [RTX 4070](https://modelfit.io/gpu/rtx-4070/) - [RTX 4070 SUPER](https://modelfit.io/gpu/rtx-4070-super/) - [RTX 5070](https://modelfit.io/gpu/rtx-5070/) - [RTX 5070 Ti](https://modelfit.io/gpu/rtx-5070-ti/) - [RTX 4070 Ti SUPER](https://modelfit.io/gpu/rtx-4070-ti-super/) - [RTX 4080 SUPER](https://modelfit.io/gpu/rtx-4080-super/) - [RTX 5080](https://modelfit.io/gpu/rtx-5080/) - [RTX 3090](https://modelfit.io/gpu/rtx-3090/) - [RTX 4090](https://modelfit.io/gpu/rtx-4090/) - [RTX 5090](https://modelfit.io/gpu/rtx-5090/) - [NVIDIA RTX PRO 6000 Blackwell](https://modelfit.io/gpu/rtx-6000-pro/) - [AMD Radeon RX 7900 XTX](https://modelfit.io/gpu/rx-7900-xtx/) - [AMD Radeon RX 7900 XT](https://modelfit.io/gpu/rx-7900-xt/) - [AMD Ryzen AI Max+ 395 (Strix Halo)](https://modelfit.io/gpu/ryzen-ai-max-395/) ## Latest Blog Posts - [Gemma 4 GPU Requirements: Every Variant (2026)](https://modelfit.io/blog/gemma-4-gpu-requirements/) - [GPT-OSS 20B VRAM Requirements (2026)](https://modelfit.io/blog/gpt-oss-20b-vram-requirements/) - [NVIDIA PAIR: Your Mac Just Became a Node in a Local AI Cluster (2026)](https://modelfit.io/blog/nvidia-pair-mac-local-ai-cluster/) ## Comparisons - [Qwen vs Llama](https://modelfit.io/compare/qwen-vs-llama/) - [Qwen vs DeepSeek](https://modelfit.io/compare/qwen-vs-deepseek/) - [Llama vs Mistral](https://modelfit.io/compare/llama-vs-mistral/) - [DeepSeek vs Llama](https://modelfit.io/compare/deepseek-vs-llama/) - [Gemma vs Phi](https://modelfit.io/compare/gemma-vs-phi/) - [Mistral vs Qwen](https://modelfit.io/compare/mistral-vs-qwen/) - [Phi vs Llama](https://modelfit.io/compare/phi-vs-llama/) - [Apple M4 vs Apple M3](https://modelfit.io/compare/m4-vs-m3-llm/) - [Apple M4 Pro vs Apple M4 Max](https://modelfit.io/compare/m4-pro-vs-m4-max-llm/) - [Apple M5 Pro vs Apple M5 Max](https://modelfit.io/compare/m5-pro-vs-m5-max-llm/) - [Mac Mini M4 vs Mac Studio M4 Max](https://modelfit.io/compare/mac-mini-vs-mac-studio-llm/) - [16 GB RAM vs 32 GB RAM](https://modelfit.io/compare/16gb-vs-32gb-ram-llm/) - [8 GB RAM vs 16 GB RAM](https://modelfit.io/compare/8gb-vs-16gb-ram-llm/) - [NVIDIA RTX 4070 (12 GB) vs Apple M4 (16-32 GB unified)](https://modelfit.io/compare/rtx-4070-vs-mac-m4-llm/) - [NVIDIA RTX 5070 (12 GB) vs NVIDIA RTX 4080 (16 GB)](https://modelfit.io/compare/rtx-5070-vs-rtx-4080-llm/) - [NVIDIA GPU (Dedicated VRAM) vs Apple Silicon (Unified Memory)](https://modelfit.io/compare/gpu-vs-apple-silicon-llm/) - [NVIDIA RTX 5070 Ti (16 GB) vs NVIDIA RTX 5080 (16 GB)](https://modelfit.io/compare/rtx-5070-ti-vs-rtx-5080-llm/) ## AI Coding Tools - [Best models for Claude Code](https://modelfit.io/tools/claude-code/) - [Best models for OpenCode](https://modelfit.io/tools/opencode/) - [Best models for OpenClaw](https://modelfit.io/tools/openclaw/) - [Best models for Aider](https://modelfit.io/tools/aider/) - [Best models for Cline](https://modelfit.io/tools/cline/) - [Best models for Roo Code](https://modelfit.io/tools/roo-code/) - [Best models for Continue.dev](https://modelfit.io/tools/continue-dev/) - [Best models for Goose](https://modelfit.io/tools/goose/) - [Best models for Open Claude Code](https://modelfit.io/tools/open-claude-code/) > Truncated to fit the 10KB llms.txt budget. Complete URL list: https://modelfit.io/sitemap.xml