llama.cpp Release b11471 Adds PLaMo FIM Token Support
The llama.cpp project released version b11471, introducing support for Fill-in-the-Middle (FIM) vocabulary tokens for PLaMo models. This update enables proper handling of code completion and text infilling tasks when running PLaMo models through llama.cpp. It expands the range of supported model features for developers relying on local LLM inference engines. The changes were added in pull request #30090 to update the tokenizer vocabulary definitions. Updated binaries for release b11471 are available across Windows, Linux, macOS, Android, and Snapdragon backends.
## BACKGROUND
PLaMo is a family of foundation language models developed by Preferred Networks (PFN), designed for English and Japanese language capabilities. Fill-in-the-Middle (FIM) is a training and inference technique that allows language models to insert text into the middle of existing context using special tokens, which is widely used in code completion toolings.