What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
A new research paper identifies that how AI models aggregate token scores across decoding steps is just as critical as the scoring functions themselves for efficient KV cache compression. By focusing on temporal aggregation and rank preservation, the authors propose a more effective method for maintaining performance when aggressively reducing memory usage during text generation.
Covered by 1 source
- AarXiv CS.AI↗Bo Zeng, Yu Zhao, Yefeng Liu, Zhihong Lu, Xuanfan Ni, Xintong WangSep 4