More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Local-AI performance on a Mac hinges on memory bandwidth, not CPU cores, chip generation or Apple’s Neural Engine. Generating each token in a 4-bit quantized 7B model means reading about 4.4 GB of weights from memory. At 50 tokens/sec you need 220 GB/s of bandwidth. Apple’s GPUs are more than capable of the math; they’re starved by the memory pipe. So tokens/sec ≈ bandwidth ÷ model size × efficiency (roughly 0.7).
A quick chart using Apple’s specs shows base-M4 (2024) delivers 19 tok/s on a 7B model, while a 2021 M1 Max hits 64 tok/s. The same M1 Pro (200 GB/s) beats the M4 on 14B models (16 vs 10 tok/s). Four generations of base-chip upgrades—M1 to M4—only raised bandwidth from 68 to 120 GB/s. But jumping tiers in one generation (base→Pro→Max) triples or quadruples it. That means any Max-tier chip outpaces every base chip, new or old, and only Max-tier machines feel fluid on 32B models (around 270 GB/s needed).
That math upends buying logic. A used 2021 M1 Max MacBook Pro with 400 GB/s and 32 GB or 64 GB of unified memory often costs as much as a new M4 MacBook Air. Yet it outperforms that Air on every local-AI workload except media engines and power efficiency. Memory size sets your model ceiling—8 GB caps you at 7B, 16 GB at 14B, 32 GB at 32B, 64 GB at 70B—so RAM and bandwidth are the specs that matter and age least. Cloud AI users won’t care; but anyone running models locally can grab used Max-tier Macs at a steep discount.
Questions about this article
No questions yet.