AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Researchers introduced AdaRoPE, a technique that allows Transformers to apply different frequency schedules and scaling to individual attention heads rather than using a uniform configuration. By adjusting positional information based on the specific requirements of each head, the method improves performance on retrieval and long-context tasks compared to standard Rotary Position Embedding implementations.
Covered by 1 source
- AarXiv CS.AI↗Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian LiJul 23