Byte Seed team found DeepSeek "pumping" reason, long context probability can drift字节 Seed 团队发现 DeepSeek“抽风”原因,长上下文可能性能漂移
IT House October 9th news, byte Seed team submitted a paper at the end of September this year on the preprint platform arXiv, talking about the phase sensitivity caused by block KV cache compression, pointing directly to the DeepSeek "pumping" reason. The study evaluated both basic and post-training versions...
IT之家 10 月 9 日消息,字节 Seed 团队今年 9 月底在预印本平台 arXiv 提交了一篇论文,谈到分块 KV 缓存压缩带来的相位敏感性, 直指 DeepSeek“抽风”原因 。 该研究评估了基础版和后训练版…
Key Points要点速览
IT House October 9th news, byte Seed team submitted a paper at the end of September this year on the preprint platform arXiv, talking about the phase sensitivity caused by block KV cache compression, pointing directly to the DeepSeek "pumping" reason. The study evaluated basic and post-trained DeepSeek-V4-Flash and DeepSeek-V4-Pro, as well as post-trained DeepSeek-V4.1-Flash. The model compresses consecutively marked windows into fewer cache entries by blocking KV cache compression to a fixed stride, reducing the memory and attention costs of long contextual reasoning. However, this compression also introduces a new position coordinate: the phase of the token, or its position relative to the boundary of the compressed window. The team found a systemic asymmetry in models using this compression: the same information is easy to retrieve at one stage, but difficult to retrieve at another. The team refers to this periodic change in retrieval performance as phase sensitivity. In large open weight models employing such compression, long context retrieval accuracy can vary by up to 40 percentage points between stages, revealing periodic weaknesses that average benchmark scores may mask. IT Home with thesis link: https://arxiv.org/abs/2609.36322
IT之家 10 月 9 日消息,字节 Seed 团队今年 9 月底在预印本平台 arXiv 提交了一篇论文,谈到分块 KV 缓存压缩带来的相位敏感性, 直指 DeepSeek“抽风”原因 。 该研究评估了基础版和后训练版的 DeepSeek-V4-Flash 和 DeepSeek-V4-Pro,以及后训练后的 DeepSeek-V4.1-Flash。 模型通过分块 KV 缓存压缩以固定步幅将连续标记的窗口压缩为更少的缓存条目,能够减少长上下文推理的内存和注意力成本。 然而,这种压缩还引入了一个新的位置坐标: Token 的相位 ,或其相对于压缩窗口边界的位置。 该团队发现使用这种压缩的模型中存在系统性不对称性: 相同的信息在一个阶段很容易检索,但在另一个阶段很难检索 。团队将这种检索性能的周期性变化称为相位敏感性(phase sensitivity)。 在采用此类压缩的大型开放权重模型中,长上下文检索准确度在各个阶段之间可能 相差高达 40 个百分点 ,从而揭示了平均基准分数可能掩盖的周期性弱点。 IT之家附论文链接: https://arxiv.org/abs/2609.36322
想马上用起来?去「AI工具」栏挑一个直接下载,或在「AI教程」里跟着图文步骤做一遍。
去 AI工具 → 看 AI教程 →阅读与点赞数据保存在你的浏览器本地,欢迎留下你的想法。
今日正能量学一点,用一点;今天种下的种子,会长成明天的能力。去免费下载专区 →广告
评论 文明发言,让讨论更有价值