Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
Researchers have published a study testing whether frontier language models can accurately report on their own internal computational states. The findings examine the extent to which these systems can identify and describe changes occurring within their architecture during processing, providing a baseline for assessing the reliability of model introspection.
Covered by 1 source
- AarXiv CS.AI↗Emilio FerraraAug 24