At 48GB, long context becomes a real workflow. The 35GB AI budget carries a 9B-14B model holding 64K-128K tokens. It digests entire codebases, contracts, or research stacks in one window. Active cooling keeps the minutes-long initial read productive. The 307 GB/s M5 Pro path cuts prompt-processing time versus smaller machines.
At roughly 96,000 words, 128K tokens swallows a short novel, a quarter of dense legal discovery, or the source of a mid-size project. Choose by what binds: a 9B at full 128K, or a 14B at 64K. Expect a real pause before the first token on huge prompts. After the read, follow-up questions answer quickly.