Qwen3.5 4B Instruct
Qwen / 4B / Q4_K_M / ~3.5 GB
Best for: Coding, Agents, Multimodal·Pop: 88/100
Perf: ~11 tok/s · first token ~1.3s
Best for coding, agents, multimodal. Strong fit for 8 GB RAM with balanced speed and quality.
The iPhone 16 Pro is the notebook in your pocket: a 4B model for capturing ideas, sketching dialogue, and unblocking a scene from wherever the idea strikes. Capture on the phone; compose on the Mac.
Creative writing on an iPhone 16 Pro is for drafts, not novels. The 6GB AI budget fits 2B-4B models that handle short stories, social posts, and scene sketches at conversational speed. The roughly 60 GB/s A18 Pro memory path (est.) keeps bursts flowing. Passive cooling only matters if you generate for many minutes at once.
The value is catching ideas wherever they land. Dictate a scene in the coffee shop, polish it on the train, and the text stays on the phone until you want it elsewhere. Keep prompts and outputs short, because the 8GB budget has no room for long context. For serious drafting, transfer the text to a Mac and let a bigger model take over.
Qwen / 4B / Q4_K_M / ~3.5 GB
Best for: Coding, Agents, Multimodal·Pop: 88/100
Perf: ~11 tok/s · first token ~1.3s
Best for coding, agents, multimodal. Strong fit for 8 GB RAM with balanced speed and quality.
Gemma / 4.5B / Q4_K_M / ~4 GB
Best for: On-device, Mobile, Chat·Pop: 82/100
Perf: ~10 tok/s · first token ~1.4s
Best for on-device, mobile, chat. Strong fit for 8 GB RAM with balanced speed and quality.
Phi / 3.8B / Q4_K_M / ~3.2 GB
Best for: Coding, Chat·Pop: 75/100
Perf: ~12 tok/s · first token ~1.3s
Best for coding, chat. Strong fit for 8 GB RAM with balanced speed and quality.
Gemma / 4B / Q4_K_M / ~3.5 GB
Best for: Chat, Coding·Pop: 81/100
Perf: ~11 tok/s · first token ~1.3s
Best for chat, coding. Strong fit for 8 GB RAM with balanced speed and quality.
Phi / 3.8B / Q4_K_M / ~3.2 GB
Best for: Coding, Chat·Pop: 64/100
Perf: ~12 tok/s · first token ~1.3s
Best for coding, chat. Strong fit for 8 GB RAM with balanced speed and quality.
Gemma / 2.3B / Q4_K_M / ~2.3 GB
Best for: IoT, Mobile, Edge·Pop: 76/100
Perf: ~20 tok/s · first token ~1.0s
Best for iot, mobile, edge. Strong fit for 8 GB RAM with balanced speed and quality.
Qwen / 2B / Q4_K_M / ~1.8 GB
Best for: Chat, Edge tasks·Pop: 75/100
Perf: ~23 tok/s · first token ~0.9s
Best for chat, edge tasks. Strong fit for 8 GB RAM with balanced speed and quality.
Llama / 3B / Q4_K_M / ~2.5 GB
Best for: Chat·Pop: 72/100
Perf: ~15 tok/s · first token ~1.1s
Best for chat. Strong fit for 8 GB RAM with balanced speed and quality.
As a thinking tool. Voice-memo a premise and have the model expand it into bullets; ask for five complications to a scene while in line for coffee; draft a character monologue on the train. The 4B class is great at idea-volume and rough sketches, exactly what mobile moments are for.
Do not draft chapters here: small-model prose plus a phone keyboard is the wrong tool twice over. Apps with iCloud-synced history make the handoff natural, and the sketch you made at lunch is waiting in context when you sit down at the Mac.
Use the ModelFit wizard to size a writing model for your exact iPhone 16 Pro RAM and context needs.
Open ModelFit Wizard