Qwen3.5 9B Instruct
Qwen / 9B / Q4_K_M / ~7 GB
Best for: Quality, Coding, Reasoning·Pop: 86/100
Perf: ~22 tok/s · first token ~0.9s
Best for quality, coding, reasoning. Strong fit for 16 GB RAM with balanced speed and quality.
A MacBook Air with a local model is a fully self-contained AI: open the lid, work with confidential material, and verify with airplane mode that nothing can leave. For client work under NDA, 16GB covers the 4B-9B class.
Privacy on a MacBook Air is a memory story, not a speed story. The 11GB AI budget on the 16GB default decides which offline models you can carry. The fanless chassis throttles only under sustained output, which confidential work rarely needs. The 153 GB/s M5 path makes short bursts feel instant. The 8GB to 32GB lineup stretches the same private stack further on bigger configs.
Turn wifi off and load a 4B or 9B model: the whole analysis happens on the machine, which is the point. A FileVault-encrypted Air running Ollama holds no cloud path by construction, a simpler story than any vendor DPA. The practical ceiling is 9B for document Q&A and drafting; 14B models fit only in the 24GB and 32GB configs. Data at rest and in use stays yours.
Qwen / 9B / Q4_K_M / ~7 GB
Best for: Quality, Coding, Reasoning·Pop: 86/100
Perf: ~22 tok/s · first token ~0.9s
Best for quality, coding, reasoning. Strong fit for 16 GB RAM with balanced speed and quality.
Qwen / 8B / Q4_K_M / ~6.5 GB
Best for: Chat, Coding·Pop: 88/100
Perf: ~25 tok/s · first token ~0.8s
Best for chat, coding. Strong fit for 16 GB RAM with balanced speed and quality.
Gemma / 12B / Q4_K_M / ~8 GB
Best for: Chat, Coding, Multimodal·Pop: 80/100
Perf: ~17 tok/s · first token ~1.0s
Best for chat, coding, multimodal. Strong fit for 16 GB RAM with balanced speed and quality.
Ornith / 9B / Q4_K_M / ~5.6 GB
Best for: Agentic coding on small machines·Pop: 76/100
Perf: ~22 tok/s · first token ~0.9s
Best for agentic coding on small machines. Strong fit for 16 GB RAM with balanced speed and quality.
Llama / 8B / Q4_K_M / ~6.5 GB
Best for: Chat, Coding·Pop: 78/100
Perf: ~25 tok/s · first token ~0.8s
Best for chat, coding. Strong fit for 16 GB RAM with balanced speed and quality.
Gemma / 12B / Q4_K_M / ~9.5 GB
Best for: Chat, Quality·Pop: 76/100
Perf: ~17 tok/s · first token ~1.0s
This model may feel memory-heavy on 16 GB RAM, but it is still listed for balanced speed and quality.
Mistral / 12B / Q4_K_M / ~9.5 GB
Best for: Chat, Translation·Pop: 78/100
Perf: ~17 tok/s · first token ~1.0s
This model may feel memory-heavy on 16 GB RAM, but it is still listed for balanced speed and quality.
Qwen / 4B / Q4_K_M / ~3.5 GB
Best for: Coding, Agents, Multimodal·Pop: 88/100
Perf: ~50 tok/s · first token ~0.6s
Best for coding, agents, multimodal. Strong fit for 16 GB RAM with balanced speed and quality.
Turn the network off. A local Ollama model behaves identically with wifi disabled. That simple test is your proof, repeatable any time, no trust required. Contracts, medical letters, financial statements: summarize and analyze them in a coffee shop without the contents existing anywhere but your SSD.
For consultants, the Air is the portable privacy story: client data processed on-site never touches your cloud accounts. The 9B class handles document Q&A and summarization; disk encryption (FileVault) closes the at-rest side of the loop.
Confirm your private workflow fits your exact MacBook Air by running the ModelFit wizard with your memory.
Open ModelFit Wizard