Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers
Researchers have found that when large language models audit one another using procedural traces, the overseer models often become overly reliant on the provided steps rather than the final output. This tendency can lead to gullibility, where an overseer ignores factual inaccuracies if the supporting rationale appears structured and logical. These findings highlight a critical vulnerability in current oversight pipelines that rely on AI-generated explanations to verify content.
Covered by 1 source
- AarXiv CS.AI↗Zihan Chen, Di Zhu, Lei Zheng, Weiling Li5d ago