← Back to Model Beat
Opinion·22h ago·all news from September 15, 2026

Refusal Reads Only a Slice of What the Model Knows: Harm-Keyed Routing and Its Exceptions Across Model Families

Researchers have identified that AI refusal mechanisms are often shallow, relying on specific activation patterns within a model's residual stream. By editing these localized directions, the safety guardrails protecting a model can be systematically bypassed. This discovery suggests that current post-training alignment techniques do not deeply integrate safety across a model's entire knowledge base. Consequently, this vulnerability highlights a fundamental limitation in how developers enforce constraints, potentially leaving models susceptible to adversarial removal of their refusal behaviors.

Covered by 1 source

Related stories

OpinionAI for Societal ImpactSep 15OpinionCognition helps Devin test its own work with GPT‑6 AstraSep 11OpinionFixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at ScaleSep 14OpinionA Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal AnalysisSep 15