LLMs respond differently to harmful prompts when AI watermarking is used
Research indicates that applying SynthID watermarks to large language models can inadvertently decrease their ability to refuse harmful prompts. This vulnerability suggests that the process of embedding digital identifiers may interfere with the safety guardrails established during model training.
Covered by 2 sources · 3 articles
- AArs Technica↗Dan Goodin4d ago
- HHacker News↗MC9954d ago
- HHacker News↗fourfire4d ago