Qwen3.6 35B-A3B
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~39 tok/s · first token ~1.5s
Best for reasoning, coding, agents. Strong fit for 48 GB RAM with balanced speed and quality.
At 32GB, a MacBook Pro crosses the line where AI prose stops needing apologies. The 14B-24B class writes with real sentence variety, holds character voice across a chapter, and takes style direction seriously.
A MacBook Pro is the laptop a serious local writer wants. The 35GB AI budget on the 48GB default fits 14B-24B models. Their prose is noticeably steadier than the 9B class. Active cooling sustains long drafting sessions at full speed. The 307 GB/s M5 Pro path keeps 30-minute generation stretches productive instead of sluggish.
The jump from 7B to 14B is where output stops feeling repetitive, and 24B-class models hold voice across long passages. Pair the model with a long context window so it remembers characters and plot threads. Draft, then use a second pass with a tighter prompt for line editing. Sensitive themes that cloud filters refuse are no problem locally.
Qwen / 35B / Q4_K_M / ~22 GB
Best for: Reasoning, Coding, Agents·Pop: 88/100
Perf: ~39 tok/s · first token ~1.5s
Best for reasoning, coding, agents. Strong fit for 48 GB RAM with balanced speed and quality.
Qwen / 35B / Q4_K_M / ~20 GB
Best for: Reasoning, Coding, Agent scenarios·Pop: 90/100
Perf: ~39 tok/s · first token ~1.5s
Best for reasoning, coding, agent scenarios. Strong fit for 48 GB RAM with balanced speed and quality.
Qwen / 27B / Q4_K_M / ~18 GB
Best for: Coding, Quality, Long context·Pop: 92/100
Perf: ~15 tok/s · first token ~1.1s
Best for coding, quality, long context. Strong fit for 48 GB RAM with balanced speed and quality.
Qwen / 30B / Q4_K_M / ~22 GB
Best for: Quality, Coding·Pop: 78/100
Perf: ~42 tok/s · first token ~1.5s
Best for quality, coding. Strong fit for 48 GB RAM with balanced speed and quality.
Gemma / 31B / Q4_K_M / ~20 GB
Best for: Quality, Coding, Multimodal·Pop: 84/100
Perf: ~13 tok/s · first token ~2.0s
Best for quality, coding, multimodal. Strong fit for 48 GB RAM with balanced speed and quality.
Gemma / 26B / Q8_0 / ~28.1 GB
Best for: Chat, Coding, Multimodal·Pop: 86/100
Perf: ~21 tok/s · first token ~0.9s
This model may feel memory-heavy on 48 GB RAM, but it is still listed for balanced speed and quality.
Gemma / 26B / Q4_K_M / ~16 GB
Best for: Chat, Coding, Multimodal·Pop: 86/100
Perf: ~39 tok/s · first token ~0.7s
Best for chat, coding, multimodal. Strong fit for 48 GB RAM with balanced speed and quality.
Qwen / 27B / Q4_K_M / ~16.5 GB
Best for: Coding, Agent, Vision, Long context·Pop: 95/100
Perf: ~15 tok/s · first token ~1.1s
Best for coding, agent, vision, long context. Strong fit for 48 GB RAM with balanced speed and quality.
Vocabulary and consistency. The jump from 9B to 14B is the most noticeable quality step in local writing models: fresher phrasing, fewer tics, dialogue that stays in character. The ~22GB budget also fits a chapter plus your style notes and a character sheet in context simultaneously.
Use system prompts as a style contract (voice, tense, banned cliches) and the 14B class actually honors it. For long projects, pin a story-bible summary at the top of context and regenerate it as the plot moves.
Use the ModelFit wizard to size a writing model for your exact MacBook Pro RAM and context needs.
Open ModelFit Wizard