llama.cpp Release b11318 Fixes Tokenizer Settings for PLaMo-2 and PLaMo-3 Models
llama.cpp release b11318 updates vocabulary handling to properly write and honor BOS (Beginning of Sequence) and EOS (End of Sequence) metadata for PLaMo-2 and PLaMo-3 models. The fix ensures that tokenizer flags such as `add_bos_token: true` and `add_eos_token: false` are correctly applied during model conversion and tokenization. Correct tokenization metadata ensures that LLM inference produces accurate responses without unexpected formatting issues or token misalignments. This patch fixes prompt handling discrepancies when executing Japanese PLaMo models via llama.cpp. Previously, `_set_vocab_plamo()` omitted BOS/EOS metadata, causing the PLAMO2 tokenization logic to ignore `add_bos` and `add_eos` parameters. Existing GGUF files that lack these metadata keys will preserve their prior behavior to maintain backward compatibility.
## BACKGROUND
llama.cpp is an open-source C/C++ inference framework designed to run large language models locally using the GGUF file format. Special tokens like BOS and EOS are critical markers used by tokenizers to signal where text generation starts and ends. PLaMo is a family of generative language models developed by Preferred Networks (PFN) with a strong emphasis on Japanese language tasks.