GitHub Copilot Adds Local AI Model Inference Support with On-Device Orchestration
Microsoft announced that GitHub Copilot will support local AI model inference by late October, introducing options for automatic or manual switching between cloud and on-device execution. Developers can run local models like Microsoft's new MAI Code 1.1 Flash or connect custom OpenAI-compatible endpoints. This update reduces dependency on cloud connectivity, providing faster response times, offline coding capabilities, and enhanced code privacy for developers. It marks a major shift toward hybrid AI orchestration in mainstream developer tools. Microsoft highlighted its MAI Code 1.1 Flash Mixture-of-Experts model, which features 137 billion total parameters but only 6.8 billion active parameters, optimized via quantization and speculative decoding. Tested on the Surface Laptop Ultra, the model achieved local decoding throughput between 40 and 63 tokens per second across prompt lengths from 2K to 256K tokens.
## BACKGROUND
Traditionally, code completion tools like GitHub Copilot relied heavily on cloud-hosted Large Language Models (LLMs) due to the high computational power required for AI inference. On-device inference relies on techniques like model quantization to run LLMs locally on consumer hardware, eliminating network latency and enabling offline use. Hybrid AI orchestration allows applications to intelligently route tasks between local hardware and cloud infrastructure based on latency, cost, and task complexity.