Qwen3.5 9B Instruct
Qwen / 9B / Q4_K_M / ~7 GB
Best for: Quality, Coding, Reasoning·Pop: 86/100
Perf: ~22 tok/s · first token ~0.9s
Best for quality, coding, reasoning. Strong fit for 16 GB RAM with balanced speed and quality.
A MacBook Air M4 with 16GB RAM runs coding models in the 4B-9B class well, with one caveat: no fan. Short completions are instant, but a 20-minute agentic session will warm the chassis and shave off speed.
Coding is the Air's most demanding realistic workload. Agent loops keep the M5 busy for minutes, and a fanless chassis gives up speed as it warms. The 153 GB/s memory path keeps short completions instant, which is where most IDE autocomplete lives. ModelFit covers Air configs from 8GB to 32GB, and the 16GB default leaves an 11GB AI budget. That fits 9B-class coders and small 14B models. Long refactors still work, just slower after the first few minutes.
Prefer 4B models for interactive editing and 9B for review sessions. Pin the small coder next to your editor, and let Ollama swap in the 9B for chat. Keep context near 8K to 16K on 16GB, because every file you add to the prompt competes with the weights. Q4 quantizations also run cooler, which keeps the chassis out of the thermal danger zone.
Qwen / 9B / Q4_K_M / ~7 GB
Best for: Quality, Coding, Reasoning·Pop: 86/100
Perf: ~22 tok/s · first token ~0.9s
Best for quality, coding, reasoning. Strong fit for 16 GB RAM with balanced speed and quality.
Qwen / 8B / Q4_K_M / ~6.5 GB
Best for: Chat, Coding·Pop: 88/100
Perf: ~25 tok/s · first token ~0.8s
Best for chat, coding. Strong fit for 16 GB RAM with balanced speed and quality.
Gemma / 12B / Q4_K_M / ~8 GB
Best for: Chat, Coding, Multimodal·Pop: 80/100
Perf: ~17 tok/s · first token ~1.0s
Best for chat, coding, multimodal. Strong fit for 16 GB RAM with balanced speed and quality.
Ornith / 9B / Q4_K_M / ~5.6 GB
Best for: Agentic coding on small machines·Pop: 76/100
Perf: ~22 tok/s · first token ~0.9s
Best for agentic coding on small machines. Strong fit for 16 GB RAM with balanced speed and quality.
Llama / 8B / Q4_K_M / ~6.5 GB
Best for: Chat, Coding·Pop: 78/100
Perf: ~25 tok/s · first token ~0.8s
Best for chat, coding. Strong fit for 16 GB RAM with balanced speed and quality.
Gemma / 12B / Q4_K_M / ~9.5 GB
Best for: Chat, Quality·Pop: 76/100
Perf: ~17 tok/s · first token ~1.0s
This model may feel memory-heavy on 16 GB RAM, but it is still listed for balanced speed and quality.
Mistral / 12B / Q4_K_M / ~9.5 GB
Best for: Chat, Translation·Pop: 78/100
Perf: ~17 tok/s · first token ~1.0s
This model may feel memory-heavy on 16 GB RAM, but it is still listed for balanced speed and quality.
Qwen / 4B / Q4_K_M / ~3.5 GB
Best for: Coding, Agents, Multimodal·Pop: 88/100
Perf: ~50 tok/s · first token ~0.6s
Best for coding, agents, multimodal. Strong fit for 16 GB RAM with balanced speed and quality.
The Air throttles under sustained load, and coding assistants are exactly that: an agent loop or a long refactor keeps the GPU busy for minutes at a stretch. Favor a 4B coder for autocomplete and quick edits, and reserve the 9B class for code review sessions where you can tolerate the slowdown after the first few minutes.
Pair the model with an editor extension like Continue.dev or Cline pointed at Ollama. Keep context windows modest (8K-16K) on 16GB: every open file you stuff into the prompt costs RAM that competes with the model weights.
Run the ModelFit wizard with your exact MacBook Air to see which coding models fit your RAM and chip.
Open ModelFit Wizard