Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families
Researchers have identified that AI refusal mechanisms are often shallow, relying on specific activation patterns within a model's residual stream. By editing these localized directions, the safety guardrails protecting a model can be systematically bypassed. This discovery suggests that current post-training alignment techniques do not deeply integrate safety across a model's entire knowledge base. Consequently, this vulnerability highlights a fundamental limitation in how developers enforce constraints, potentially leaving models susceptible to adversarial removal of their refusal behaviors.
Covered by 1 source
- AarXiv CS.AI↗Orion Reblitz-Richardson22h ago