DeepSeek logo

DeepSeek Models: Cloud V4, Local R1 Reasoning

DeepSeek's V4 generation is a cloud story. V4.1 Flash (released 2026-09-10, MIT) is a multimodal image + text MoE with a 552B backbone and a 196B Engram conditional memory that activates 8B on prefill and 16B on decode, with 1M context. V4 Flash 0731 (284B total / 13B active) and V4 Pro (1.6T total / 49B active) are API models, and none of the three has a registry-verified Ollama tag. The local DeepSeek line is R1: the 7B, 14B and Llama 70B distills run in Q4_K_M via Ollama starting at 8GB RAM, and the full R1 671B needs a 512GB Mac Studio. V4 Flash 0731 does get community 2-bit dynamic quants that fit 128GB Macs through MLX, llama.cpp or the ds4 engine.

DeepSeek AI6 local models+ 6 API
DEVELOPER
DeepSeek AI
MODELS
6
SIZE RANGE
7B–671B
RAM RANGE
8–512 GB
Bar chart: maximum local LLM size by memory tier. 8 GB runs up to 9B, 12 GB runs up to 12B, 16 GB runs up to 14B, 24 GB runs up to 27B, 32 GB runs up to 35B, 36 GB runs up to 35B, 48 GB runs up to 35B, 64 GB runs up to 70B, 72 GB runs up to 70B, 96 GB runs up to 70B, 128 GB runs up to 70B, 192 GB runs up to 70B, 256 GB runs up to 70B, 512 GB runs up to 405B. Data from ModelFit's own catalog.
Key Features
V4.1 Flash: multimodal 552B MoE, API only
R1 distills 7B-70B run locally in Ollama
Full R1 671B needs a 512GB Mac Studio
V4 Flash 0731 quants fit 128GB Macs
Reasoning-first: strong for coding and debugging

All DeepSeek Models

ModelSizeQuantVRAMMin RAMBest ForQualityOllama
DeepSeek-R1 Distill Qwen 7B7BQ4_K_M5.5 GB8 GBReasoning, Coding
80
DeepSeek-R1 Distill Qwen 7B (Q8)7BQ8_07.5 GB16 GBReasoning, Coding
82
DeepSeek-R1 Distill Qwen 14B14BQ4_K_M11 GB16 GBReasoning, Quality
86
DeepSeek-R1 Distill Qwen 14B (Q8)14BQ8_014.6 GB24 GBReasoning, Quality
88
DeepSeek-R1 Distill Llama 70B70BQ4_K_M42 GB64 GBReasoning, Quality
85
DeepSeek-R1 671B671BQ4_K_M380 GB512 GBReasoning, Coding
96

Cloud / API-only models

Weights are open but far beyond consumer hardware — these run as APIs. Each page lists the local alternatives.

ModelSizeAccessBest For
DeepSeek V4 Flash 0731284BAPI onlyReasoning, Coding, Long context
DeepSeek V4.1 Flash552BAPI onlyAgentic coding, Reasoning, Long context, Vision
DeepSeek-V3671BAPI onlyQuality, Coding
DeepSeek-V3-0324671BAPI onlyQuality, Coding
DeepSeek-R1671BAPI onlyReasoning, Quality
DeepSeek V4 Pro1600BAPI onlyFrontier reasoning, Coding

Device Compatibility

Which DeepSeek models can run on each device class, based on minimum RAM requirements.

ModeliPhoneAirProStudioMini
DeepSeek-R1 Distill Qwen 7B (7B)PossiblePossibleExcellentExcellentExcellent
DeepSeek-R1 Distill Qwen 7B (Q8) (7B)NoPossiblePossibleExcellentPossible
DeepSeek-R1 Distill Qwen 14B (14B)NoPossiblePossibleExcellentPossible
DeepSeek-R1 Distill Qwen 14B (Q8) (14B)NoPossiblePossiblePossiblePossible
DeepSeek-R1 Distill Llama 70B (70B)NoNoPossiblePossiblePossible
DeepSeek-R1 671B (671B)NoNoNoPossibleNo

RAM Requirements

DeepSeek-R1 Distill Qwen 7B
5.5 GB · min 8 GB
DeepSeek-R1 Distill Qwen 7B (Q8)
7.5 GB · min 16 GB
DeepSeek-R1 Distill Qwen 14B
11 GB · min 16 GB
DeepSeek-R1 Distill Qwen 14B (Q8)
14.6 GB · min 24 GB
DeepSeek-R1 Distill Llama 70B
42 GB · min 64 GB
DeepSeek-R1 671B
380 GB · min 512 GB

Frequently Asked Questions

Can I run DeepSeek V4.1 Flash locally?
No. V4.1 Flash is a cloud-only multimodal MoE with no registry-verified Ollama tag, and the same holds for V4 Pro. The local DeepSeek line is the R1 reasoning distills; V4 Flash 0731 is the only current V4 member that runs locally, and only through community 2-bit quants on 128GB Macs.
What is the difference between V4.1 Flash, V4 Flash 0731 and V4 Pro?
V4.1 Flash (2026-09-10) is MIT-licensed and multimodal: a 552B backbone plus 196B Engram conditional memory, with 8B active on prefill and 16B on decode, and 1M context. V4 Flash 0731 is a 284B total / 13B active MoE with 1M context. V4 Pro is the 1.6T total / 49B active cloud model. All three are API-only from DeepSeek.
What Ollama commands run DeepSeek R1 locally?
Run `ollama run deepseek-r1:7b` on 8GB RAM (about 5.5GB load), `ollama run deepseek-r1:14b` on 16GB (about 11GB), or `ollama run deepseek-r1:70b` on 64GB (about 42GB). The full `ollama run deepseek-r1:671b-q4_K_M` needs a 512GB machine and loads around 380GB.
Can the newest DeepSeek flagship run on consumer hardware?
No. V4.1 Flash is cloud-only with no registry-verified build, and V4 Pro targets the API. The closest thing to local is V4 Flash 0731: 128GB-Mac owners can run it through community 2-bit dynamic quants via MLX, llama.cpp or the ds4 engine, but that is not an official Ollama tag.
Is DeepSeek R1 good for coding?
Yes. The R1 distills are reasoning-first, which helps with debugging, architecture and multi-step coding tasks. The 14B and 70B versions are the strongest local options in the family; the 7B stays fast enough for laptops.

Related Model Families

Getting Started