We’re putting too much faith in AI’s ability to say no
Researchers are raising concerns that current safety measures designed to make AI models decline harmful requests are insufficient and easily bypassed. These findings highlight a growing gap between public expectations for reliable AI refusal and the technical reality of how easily modern systems can be manipulated into ignoring their programmed constraints.
Covered by 1 source
- MMIT Technology Review↗Arthur Holland Michel8h ago