Running Large MoE LLMs on 64GB Apple Silicon Macs Using oMLX
A Reddit user demonstrated running a 100GB Qwen Mixture-of-Experts (MoE) model on an M3 Max Mac with 64GB of unified memory using the latest release of oMLX. The setup achieved a generation speed of 15.8 tokens per second and a prompt processing speed of 140 tokens per second with context windows up to 130k tokens. This achievement highlights how software optimizations like expert offloading and multi-token prediction make it possible to run large enterprise-grade AI models locally on consumer hardware. It significantly lowers memory barriers for developers and local AI enthusiasts who previously required expensive multi-GPU setups. The setup allocated 58GB of RAM to GPU memory using oMLX features like 50% resident MoE expert offload, aggressive memory guards, and Lightning Multi-Token Prediction (MTP). The session demonstrated high efficiency, maintaining a 90.9% prompt cache hit rate across more than 320,000 prefill tokens.
## BACKGROUND
oMLX is a local inference server designed for Apple Silicon Macs that builds on Apple's native MLX framework to optimize local LLM execution using unified memory architecture. Mixture-of-Experts (MoE) models use sparsely activated subnetworks called experts, allowing runtime tools to offload inactive experts to system memory or disk so giant models can run on limited hardware.