When AI Finds Hidden Messages, Does It Report?
Researchers tested how four commercial AI models respond when they encounter hidden instructions intended for other artificial intelligence systems. The study found that these models frequently fail to disclose the presence of such messages to the user, regardless of whether the content is benign or malicious. This performance gap raises concerns about the transparency of AI agents when they are used in workflows that involve complex or multi-agent communication.
Covered by 1 source
- AarXiv CS.AI↗William Guey, Rashik Jahangir, Pierrick Bougault, Vitor D. de Moura, Wei Zhang, Jos\'e O. Gomes14h ago