High-Speed Local LLM Inference on iPadOS via External AMD GPU Port
LemonSeed Studio demonstrated running the Qwen3.8-27B language model at 158.5 tokens/second on an iPad connected to an external AMD Radeon AI PRO R9700 GPU over Thunderbolt. This was achieved by porting upstream Linux kernel drivers (amdgpu and amdkfd) to iPadOS as a custom PCIDriverKit extension alongside the high-performance LemonSeed graph engine. This engineering feat breaks conventional platform boundaries by unlocking desktop-class eGPU hardware acceleration on iPadOS for demanding AI tasks. It demonstrates how combining custom driver extensions with speculative decoding techniques like DFlash2 can deliver near-native desktop LLM inference speeds on mobile operating systems. The LemonSeed Engine (LSE 0.5.8) records model forward passes as computational graphs, fuses operators, and benchmarks kernel layouts on-device to select the fastest execution path. Enabling DFlash2 draft tree speculative decoding boosted inference performance on iPadOS from a baseline of 31.2 tok/s up to 158.5 tok/s, matching macOS performance and maintaining over 1,600 tok/s prefill speeds at 4K context.
## BACKGROUND
Apple's PCIDriverKit framework allows developers to create custom user-space drivers to control PCI Express hardware connected over Thunderbolt. Speculative decoding accelerates LLM generation by employing a smaller draft model—such as block diffusion draft models like DFlash2—to quickly generate candidate tokens that the larger target model verifies in parallel within a single forward pass.