Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States
Researchers have identified that diagnostic evidence from language model internal states does not uniformly support all causal claims, as its relevance depends on the specific populations and pathways being analyzed. This finding demonstrates that identifying causal relationships within model architectures requires precise alignment between diagnostic metrics and the specific variables under investigation.
Covered by 1 source
- AarXiv CS.AI↗Weiyi Kong, Zhuoran LiJul 30