How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Researchers have proposed new metrics to evaluate the consistency and specificity of language model circuits, aiming to move beyond current assessments based solely on necessity and sufficiency. This development addresses a limitation in mechanistic interpretability, where current methods often fail to accurately identify which specific subgraphs truly drive particular model behaviors.
Covered by 1 source
- AarXiv CS.AI↗Michael Li, Nishant SubramaniSep 4