NeFut Logo NeFut
Admin Login

[CS.AI] Revolutionary KV Cache Merging: Selective Merging and Attention Compensation

Published at: 2026-07-21 22:00 Last updated: 2026-07-22 01:01
#algorithm #AI #optimization

Abstract

Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major bottleneck. Recent compression methods mitigate this cost via token merging; however, these approaches often rely on indiscriminate aggregation, which degrades representations and introduces attention sag, a mismatch where merged tokens receive the same softmax mass as individual tokens despite encoding multiple inputs.

We propose a training-free, dual-component framework for KV cache compression that addresses these limitations. First, a soft cosine gate adaptively modulates merging decisions based on value-vector similarity, suppressing or discarding dissimilar tokens to preserve semantic fidelity. Second, we introduce an attention-ratio compensation mechanism that applies a decoding-time logit bias derived from prefill attention statistics, correcting the softmax imbalance induced by merging.

Evaluated on LongBench (16 English datasets) while retaining only 25% of the KV cache, our framework achieves strong compressed performance against representative one-shot baselines. It is especially robust on the evaluated grouped-query attention (GQA) models, maintaining near-lossless generation quality. Furthermore, the method outperforms the full-cache baseline on complex multi-document QA tasks and delivers a 3.3x decoding speedup at 100k tokens.

Blogger's Review: This research effectively addresses key issues in KV cache compression through selective merging and attention compensation mechanisms, showcasing its potential applications in large-scale text generation tasks. A noteworthy advancement in the field.

Original Source: https://arxiv.org/abs/2607.16213

[h] Back to Home