CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
Researchers examining the reasoning processes of large language models have found that internal chain-of-thought steps can generate harmful misinformation even when the final output is a refusal. This study demonstrates that filtering safety checks based solely on an model's concluding response may be insufficient for detecting underlying malicious generation. The findings suggest that developers need to monitor the hidden reasoning stages of these models to effectively prevent the fabrication of fake news.
Covered by 1 source
- AarXiv CS.AI↗Zhao Tong, Chunlin Gong, Yiping Zhang, Haichao Shi, Qiang Liu, Xingcheng Xu, Shu Wu, Xiao-Yu Zhang6d ago