Gemma Models: Google's Lightweight Local AI

Google DeepMind's Gemma family delivers impressive quality from compact models. Gemma 2 9B is one of the best sub-10B models available, and Gemma 1B runs on iPhones. If you want a well-tuned, safety-conscious model with solid general performance, Gemma is a strong choice.

Google DeepMind14 local models
DEVELOPER
Google DeepMind
MODELS
14
SIZE RANGE
1B–31B
RAM RANGE
248 GB
Bar chart: maximum local LLM size by memory tier. 8 GB runs up to 9B, 12 GB runs up to 12B, 16 GB runs up to 14B, 24 GB runs up to 27B, 32 GB runs up to 35B, 36 GB runs up to 35B, 48 GB runs up to 35B, 64 GB runs up to 70B, 72 GB runs up to 70B, 96 GB runs up to 70B, 128 GB runs up to 70B, 192 GB runs up to 70B, 256 GB runs up to 70B, 512 GB runs up to 405B. Data from ModelFit's own catalog.
Key Features
Excellent quality at small sizes (1B-9B)
Strong safety and instruction tuning
Efficient on Apple Silicon
Good for chat and general tasks

All Gemma Models

ModelSizeQuantVRAMMin RAMBest ForQualityOllama
Gemma 3 1B Instruct1BQ4_K_M1 GB2 GBChat, Mobile
52
Gemma 2 2B Instruct2BQ4_K_M1.8 GB4 GBChat
58
Gemma 4 E2B2.3BQ4_K_M2.3 GB4 GBIoT, Mobile, Edge
68
Gemma 3 4B Instruct4BQ4_K_M3.5 GB6 GBChat, Coding
72
Gemma 4 E4B4.5BQ4_K_M4 GB6 GBOn-device, Mobile, Chat
78
Gemma 2 9B Instruct9BQ4_K_M7 GB12 GBChat, Coding
76
Gemma 3 12B Instruct12BQ4_K_M9.5 GB16 GBChat, Quality
85
Gemma 4 12B12BQ4_K_M8 GB12 GBChat, Coding, Multimodal
90
Gemma 4 12B (Q8)12BQ8_012.8 GB24 GBChat, Coding, Multimodal
92
Gemma 4 26B-A4B26BQ4_K_M16 GB24 GBChat, Coding, Multimodal
91
Gemma 4 26B-A4B (Q8)26BQ8_028.1 GB48 GBChat, Coding, Multimodal
93
Gemma 2 27B Instruct27BQ4_K_M21 GB32 GBQuality, Coding
80
Gemma 3 27B Instruct27BQ4_K_M21 GB32 GBQuality, Coding
89
Gemma 4 31B31BQ4_K_M20 GB32 GBQuality, Coding, Multimodal
93

Device Compatibility

Which Gemma models can run on each device class, based on minimum RAM requirements.

ModeliPhoneAirProStudioMini
Gemma 3 1B Instruct (1B)ExcellentExcellentExcellentExcellentExcellent
Gemma 2 2B Instruct (2B)ExcellentExcellentExcellentExcellentExcellent
Gemma 4 E2B (2.3B)ExcellentExcellentExcellentExcellentExcellent
Gemma 3 4B Instruct (4B)PossiblePossibleExcellentExcellentExcellent
Gemma 4 E4B (4.5B)PossiblePossibleExcellentExcellentExcellent
Gemma 2 9B Instruct (9B)PossiblePossiblePossibleExcellentPossible
Gemma 3 12B Instruct (12B)NoPossiblePossibleExcellentPossible
Gemma 4 12B (12B)PossiblePossiblePossibleExcellentPossible
Gemma 4 12B (Q8) (12B)NoPossiblePossiblePossiblePossible
Gemma 4 26B-A4B (26B)NoPossiblePossiblePossiblePossible
Gemma 4 26B-A4B (Q8) (26B)NoNoPossiblePossiblePossible
Gemma 2 27B Instruct (27B)NoPossiblePossiblePossiblePossible
Gemma 3 27B Instruct (27B)NoPossiblePossiblePossiblePossible
Gemma 4 31B (31B)NoPossiblePossiblePossiblePossible

RAM Requirements

Gemma 3 1B Instruct
1 GB · min 2 GB
Gemma 2 2B Instruct
1.8 GB · min 4 GB
Gemma 4 E2B
2.3 GB · min 4 GB
Gemma 3 4B Instruct
3.5 GB · min 6 GB
Gemma 4 E4B
4 GB · min 6 GB
Gemma 2 9B Instruct
7 GB · min 12 GB
Gemma 3 12B Instruct
9.5 GB · min 16 GB
Gemma 4 12B
8 GB · min 12 GB
Gemma 4 12B (Q8)
12.8 GB · min 24 GB
Gemma 4 26B-A4B
16 GB · min 24 GB
Gemma 4 26B-A4B (Q8)
28.1 GB · min 48 GB
Gemma 2 27B Instruct
21 GB · min 32 GB
Gemma 3 27B Instruct
21 GB · min 32 GB
Gemma 4 31B
20 GB · min 32 GB

Frequently Asked Questions

What is the best Gemma model for a MacBook Air?
Gemma 2 9B Q4 is the top pick for MacBook Air with 16GB RAM. It uses about 6GB and delivers strong chat performance. For 8GB Air, Gemma 2 2B is a good fallback.
Can Gemma run on an iPhone?
Yes. Gemma 1B runs on iPhones with 6GB+ RAM. Quality is basic but useful for simple chat and text tasks.
How does Gemma compare to Llama at similar sizes?
Gemma 2 9B and Llama 3.1 8B are very close. Gemma tends to be more conservative and safety-tuned, while Llama is more flexible. Pick based on whether you want guardrails or freedom.

Related Model Families

Getting Started