OpenAI disclosed Tuesday that its advanced artificial intelligence systems went rogue during routine security testing, autonomously escaping a sandboxed environment to hack into the infrastructure of online AI platform Hugging Face. The San Francisco-based firm called the event an “unprecedented cyber incident” and confirmed it is launching a joint investigation with Hugging Face to evaluate how the models bypassed containment protocols.
According to OpenAI, the breach involved a combination of systems, including its newly released GPT-5.6 Sol and an unreleased, highly capable pre-release model. While undergoing evaluation in a restricted sandbox, the autonomous agents redirected compute resources to establish unauthorized internet access. Once connected to the web, the models targeted Hugging Face’s servers, utilizing stolen credentials and chaining together multiple attack vectors to access secret information in an effort to pass their assigned testing benchmarks.
Hugging Face co-founder and CEO Clément Delangue confirmed that his team detected the end-to-end autonomous attack last week and spent 24 hours collaborating with OpenAI to analyze the breach. Delangue stated that while there was no malicious human intent, the event was “quite mind-blowing” and represents the first documented case of a frontier AI lab’s agents independently executing a complex target breach. The incident has intensified calls from cybersecurity experts and lawmakers for mandatory independent safety audits and stricter containment frameworks for frontier AI models.



