ByteDance Seed Team Discovers Phase Sensitivity Causing Long-Context Performance Drift in LLMs
ByteDance's Seed team published research on arXiv revealing that chunked key-value (KV) cache compression introduces "phase sensitivity" in long-context LLMs like DeepSeek. This phenomenon creates periodic weak spots based on token alignment within compression window boundaries, causing information retrieval accuracy to fluctuate by up to 40 percentage points. While chunked KV cache compression drastically reduces memory and attention computation costs for long inputs, phase sensitivity shows that traditional benchmark averages can mask severe retrieval failures. This insight is essential for building more reliable memory compression techniques and realistic evaluation standards for long-context models. The researchers evaluated multiple base and post-trained DeepSeek model variants, observing a systematic asymmetry where identical information is easy to retrieve at one phase position but significantly harder at another. This confirms that a token's relative phase inside a compression window acts as an unintended position coordinate that impacts recall.
## BACKGROUND
In Transformer-based large language models, the KV cache saves key and value tensors from previous tokens so the model does not have to recompute them during step-by-step generation. However, processing long contexts requires massive memory to store these tensors, prompting researchers to use chunked compression techniques that merge sequential token windows into fewer cache entries.