Click any tag below to further narrow down your results
Links
Apple released updated Mac mini and Mac Studio with new M6 and M5 Ultra chips, positioning them as machines for running large language models locally. The hardware refresh is purely specs-focused, but Apple's marketing now emphasizes AI workloads after macOS improvements last year made multi-Mac setups viable for distributed inference.
- Apple shipped macOS 26.2 in December with Thunderbolt 5 support for low-latency distributed AI inference, which triggered developer interest in daisy-chaining multiple Macs for larger models
- The M6 is Apple's first 2nm chip for Macs, and the M5 Ultra is now the most powerful option in the lineup, especially for AI tasks
- Developers and researchers are using stacked Mac minis and Studios as an alternative to expensive Nvidia GPU hardware for running local LLMs that exceed single-device capacity
This article argues that local-AI performance on Macs depends on memory bandwidth, not CPU cores, GPU cores, or the Neural Engine. Using a simple formula (bandwidth ÷ model size × efficiency), it shows a 2021 M1 Max outperforms a 2024 M4 base chip by over 3× on a 7B model. It recommends buying used Max-tier machines and highlights lineup quirks like the M3 Pro’s bandwidth regression.
- Memory bandwidth, not CPU/GPU/Neural Engine specs, determines local-AI token generation speed on Macs
- A 2021 M1 Max (400 GB/s) hits ~64 tok/s on a 7B model vs. ~19 tok/s on a base 2024 M4 (120 GB/s) — over 3× faster despite being three years older
- Tier jumps (base→Pro→Max) matter far more than generational upgrades: four generations of base chips only went from 68 to 120 GB/s, while switching tiers can triple or quadruple bandwidth
- Used Max-tier MacBook Pros often cost the same as a new M4 MacBook Air but outperform it on every local-AI task except power efficiency and media engines, making RAM/bandwidth the specs to prioritize when buying used
The author swaps ChatGPT Plus, Cursor and Midjourney for local AI on a 14″ MacBook Pro M5 Max. Two setups failed; a third ran locally by day nine and convinced him to re-subscribe.
- CUDA is irrelevant on Apple Silicon, yet the author describes "CUDA fallbacks" as part of the local setup overhead—an inconsistency suggesting the account may be unreliable
- Local LLMs (Qwen, Mistral) hit memory bottlenecks and slow inference on long prompts, even on an M5 Max
- Local Stable Diffusion couldn't match Midjourney's compositional consistency and prompt refinement, driving a resubscription
- Cursor's extension ecosystem proved hard to replicate locally, undermining the coding-assistant replacement