Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
Researchers have introduced a method called Tripwire designed to defend large language models against jailbreak attacks by targeting specific neurons associated with harmful content. Unlike previous intervention techniques that often degraded a model's overall performance, this approach aims to maintain utility while effectively triggering safety refusals.
Covered by 1 source
- AarXiv CS.AI↗Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun5d ago