Tethering an iPhone to a MacBook Pro Boosts LLM Prefill Speed by 44%
Developer "u/StayLameBro" created an open-source tool called "backburner" that offloads transformer layers of the Qwen3.8-27B model from a MacBook Pro (M4 Pro) to a tethered iPhone 17 Pro Max over USB-C. Using pipeline parallelism across both devices, prompt prefill speed increased by up to 44% at a 16K context length compared to running on the Mac alone. This experiment demonstrates a novel edge AI approach for bypassing memory and compute limitations on laptops by pooling hardware resources with mobile devices. It highlights how high-speed wired connections like USB-C can enable multi-device distributed LLM inference in personal computing environments. In this pipeline, the MacBook Pro processes layers 1–40 of each 256-token batch and streams activation data to the iPhone, which runs layers 41–64 using its A19 Pro GPU and Neural Engine. A key limitation is that this setup only accelerates the compute-heavy prompt prefill phase, while text generation (the decode phase) still runs entirely on the Mac.
## BACKGROUND
LLM inference occurs in two distinct phases: prefill, where the prompt is processed in parallel to compute initial network states, and decode, where output tokens are generated sequentially one by one. Layer offloading splits a transformer model's layers across multiple processors or hardware devices to relieve memory pressure and increase processing throughput.