← Back to Model Beat
Policy·Nov 21·all news from November 21, 2025

From shortcuts to sabotage: natural emergent misalignment from reward hacking - Anthropic

From shortcuts to sabotage: natural emergent misalignment from reward hacking Anthropic

Covered by 1 source

Related stories

PolicyGetty Images v Stability AI: A landmark judgment reinforcing the need for the UK government to amend its copyright laws - Wolters KluwerNov 20PolicyStrengthening our safety ecosystem with external testingNov 19PolicyAnthropic partners with Rwandan Government and ALX to bring AI education to hundreds of thousands of learners across Africa - AnthropicNov 18PolicyMitigating the risk of prompt injections in browser use - AnthropicNov 24 · 2 sources