~/LOCAL AI/running-large-moe-llms-on-64gb-apple-silicon-macs-using-omlx

Running Large MoE LLMs on 64GB Apple Silicon Macs Using oMLX

A Reddit user demonstrated running a 100GB Qwen Mixture-of-Experts (MoE) model on an M3 Max Mac with 64GB of unified memory using the latest release of oMLX. The setup achieved a generation speed of 15.8 tokens per second and a prompt processing speed of 140 tokens per second with context windows up to 130k tokens. This achievement highlights how software optimizations like expert offloading and multi-token prediction make it possible to run large enterprise-grade AI models locally on consumer hardware. It significantly lowers memory barriers for developers and local AI enthusiasts who previously required expensive multi-GPU setups. The setup allocated 58GB of RAM to GPU memory using oMLX features like 50% resident MoE expert offload, aggressive memory guards, and Lightning Multi-Token Prediction (MTP). The session demonstrated high efficiency, maintaining a 90.9% prompt cache hit rate across more than 320,000 prefill tokens.

## BACKGROUND

oMLX is a local inference server designed for Apple Silicon Macs that builds on Apple's native MLX framework to optimize local LLM execution using unified memory architecture. Mixture-of-Experts (MoE) models use sparsely activated subnetworks called experts, allowing runtime tools to offload inactive experts to system memory or disk so giant models can run on limited hardware.

## REFERENCES

## KEYWORDS

#Local AI#Apple Silicon#LLM Inference#MLX#MoE

$ subscribe --daily

Running Large MoE LLMs on 64GB Apple Silicon Macs Using oMLX | Daily News