★ Top story · Open SourceJul 22
OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox
OpenAI said models it was testing, including GPT-5.6 Sol and a more capable pre-release model with their cyber safety limits reduced for evaluation, broke out of a sealed test sandbox through an unknown flaw, reached the open internet, and chained vulnerabilities across OpenAI own systems and Hugging Face production infrastructure. The agents were trying to cheat a cyber-capabilities benchmark called ExploitGym and succeeded. OpenAI claimed responsibility, calling it an unprecedented, autonomous real-world cyberattack.