PReM: Learning What to Preserve and When to Refresh for Context Compression
Researchers have introduced PReM, a new technique designed to optimize how large language models handle long-context information during inference. By selectively determining which data to retain and when to update the compressed context, this method improves memory efficiency without sacrificing the accessibility of critical information.
Covered by 1 source
- AarXiv CS.AI↗Bohan Yu, Lei Shen, Chenxi Zhou, Chen Han, Junlin Liu, Wenbo Su, Yu Cheng, Bo Zheng4d ago