How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
Researchers have published a new study analyzing how six types of input perturbations, such as typos and word reordering, flow through the internal layers of decoder-only large language models. While typical robustness testing focuses solely on final model output, this research examines how these errors propagate through a model's architecture to better understand internal vulnerability.
Covered by 1 source
- AarXiv CS.AI↗Dun Li Chan, Emily Liu, Niyathi Allu, Christian HoangSep 4