The M6 and M5 Ultra are Apple's first chips built for linked together Mac minis and Mac Studios, the configuration developers use to run large AI models too big for one machine.
The M6 and M5 Ultra are Apple's first chips built for daisy-chained Mac minis and Mac Studios, the configuration developers use to run large AI models too big for one machine.
On a desk in a research lab, a developer connects two Mac minis with a single cable. The model that wouldn't fit on one machine now runs across both. The setup is unglamorous: Thunderbolt 5, an open-source toolkit, and the kind of small silver box many readers already know. The configuration doesn't appear in any Apple advertisement. Until this week, it wasn't a product story at all.
Apple refreshed the Mac Studio and Mac mini on Tuesday and introduced two new chips: the M6, the first 2-nanometer part in the M-series line, and the M5 Ultra, the new top-end Mac silicon. On the surface, the announcement is a specs bump. Ars Technica's Andrew Cunningham called it a hardware refresh without major new features, noting that Apple is "leaning hard into the local-AI-inference and dev use cases that weren't a design goal for earlier generations."
Apple's current desktops were not built for the workload they're being asked to run. Hobbyists, professional developers, and researchers began linking Mac minis and Mac Studios to run large language models too big for any single consumer machine. The company did not plan that workaround. It is now selling chips for it.
The mechanism has three pieces.
Unified memory. The CPU and GPU in an Apple silicon Mac share a single pool of high-bandwidth memory, which lets them work on the same model without shuttling data back and forth. That design choice is what makes a Mac mini a credible local-AI box instead of a slow one.
An open-source array framework that takes advantage of that shared memory. MLX is the software layer that lets developers write AI code that runs natively on M-series chips.
Thunderbolt 5. A very fast wired connection standard, fast enough to link multiple Macs with low enough latency to spread a model across them. Until macOS 26.2 shipped in December 2025, that low-latency link was not officially part of the platform.
Apple's macOS 26.2 release notes, as quoted by Ars Technica, describe the change plainly: the update enabled "low-latency communication between Thunderbolt 5 hosts for use cases including distributed AI inference using MLX." That sentence is the pivot. Before December 2025, daisy-chaining Mac minis to run a single model was a hack. Now it is an officially supported configuration.
The M6 is Apple's first 2nm part, a chip-manufacturing process that lets the company pack more transistors onto a single die. The M5 Ultra carries more CPU and GPU cores and the memory bandwidth that AI workloads want. Neither ships with a marketing line like "for distributed inference." Apple is letting the silicon do the endorsing.
Apple is selling consumer-priced hardware into a narrow slice of local inference work: development, research, and the kind of model running that data-center GPUs are overkill for. A four-Mac-mini cluster at retail price is still a small fraction of a single high-end Nvidia workstation, and that gap is the reason the workaround exists.
The Ars Technica piece frames Apple as "positioning this as an alternative to ultra-beefy specialized hardware featuring Nvidia GPUs." That is accurate, with one important limit: the alternative is for the work a four-Mac-mini cluster can do, not for the work a 72-GPU rack does. Apple's bet is that the first category is bigger than the second looks from a data-center vantage point.
The next signal to watch is whether Apple keeps extending the software support. macOS 26.2 was the first official nod. The new chips are the second. A third step, whether better first-party tooling, distributed-inference primitives, or a Mac Pro built for the same workload, would be the moment the workaround stops being a workaround.
Until then, the new Mac Studio and Mac mini are the first Apple desktops built to a specification its users wrote. The spec sheet didn't exist in Cupertino. It existed on the desks of researchers who plugged two small silver boxes together and found it worked.