llama.cpp Release b11439 Refactors Selective Expert Copying in ggml
Open-source project llama.cpp released build b11439, refactoring selective expert copying logic to user code in the underlying ggml library. Additionally, the release enrolls DeepSeek-V2 into the automated test suite for selective expert copy functionality. This refactoring improves modularity and maintainability in ggml's backend API, making it easier to handle Mixture-of-Experts (MoE) architectures. Incorporating models like DeepSeek-V2 into test suites ensures better stability for running sparse MoE inference across supported hardware backends. The release updates backend interfaces and commentary in `ggml-backend.h` to cleanly expose expert copying operations to external callers. Automated tests were modified to enroll two models, specifically using DeepSeek2 to validate selective expert copying functionality.
## BACKGROUND
llama.cpp is a popular open-source C/C++ framework for running LLMs efficiently on consumer devices using its underlying tensor library, ggml. Mixture-of-Experts (MoE) architectures, such as DeepSeek-V2, route tokens to specialized sub-networks called experts, which requires customized memory management and backend copying mechanisms.