The 24GB M4 Pro MacBook Pro runs the lower 40 layers (stacked model stages) of the 64 layer Qwen3.8 27B; the iPhone 17 Pro Max's A19 Pro GPU runs the upper 24.
A 24GB MacBook Pro cannot comfortably prefill a 27-billion-parameter open AI model at long context. Reddit user u/StayLameBro sidestepped the limit by splitting Qwen3.8-27B, at IQ4_XS quantization, across an M4 Pro MacBook Pro and an iPhone 17 Pro Max tethered over a 10 Gb/s USB-C cable. The laptop ran layers 1 through 40 of each 256-token batch; the iPhone's A19 Pro GPU ran layers 41 through 64 while the Mac started the next batch.
Measured prefill speedups, end-to-end: 132 to 177 tokens per second at 8K context (+35%), 109 to 157 at 16K (+44%), and 101 to 130 at 32K (+29%), per the developer's llama.cpp fork. At 140K context, the iPhone's Neural Engine compresses stale context into the model, dropping per-token write time from 279 ms to 176 ms and extending usable context from 64K to 196K–229K.
The result has been independently reproduced. An independent fork by soloptimizer re-measured the 16K jump at +44% with 31% less waiting, and jazir555's LLMDroid ports the concept to an Android phone and Windows PC over USB-C. A third-party test on Apple's upcoming MacBook Neo and iPhone Air reported first-try success. Output stayed token-identical with and without the phone.
The catch: the iPhone accelerates prefill, not text generation, and only above a 64K context window. Below that, it sits out the decode path. The setup also needs a wired 10 Gb/s USB-C cable and the custom open-source fork, not an Apple-supported feature. Latent.space's AINews framed the work as a tinkerer's hack, not a product roadmap.