llama.cpp Release b11435 Fixes CPU Data Race and Refactors Sequence Handling
llama.cpp release b11435 resolves a data race condition on CPU backends occurring during k-pool scatter operations on shared sequences. It also refactors sequence copying (seq_cp) to enforce whole-sequence operations and removes the legacy k-pool cache_safe mode. Fixing this race condition enhances stability and execution reliability when running multi-sequence inference on CPU architectures. Additionally, simplifying sequence state copying removes redundant staleness checks and overhead during context management. Instead of re-pooling shared cells across scatter entries, each shared k-pool rep is now marked only once per micro-batch. Furthermore, partial sequence copies are now explicitly rejected in hybrid memory, allowing the removal of complex stale-all workarounds across seq_rm, state_read, and state_drop.
## BACKGROUND
llama.cpp is a popular open-source inference engine designed to execute Large Language Models efficiently across various hardware backends. Sequence management and context sharing allow llama.cpp to handle parallel prompt processing and hybrid recurrent state architectures effectively.