OpenAI, Face
Digest more
OpenAI's rogue models used publicly exposed credentials across "four accounts on four services" to help facilitate the Hugging Face breach.
In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
Anthropic's Claude Mythos Preview has exposed flaws in core encryption standards just one week after OpenAI models escaped sandbox containment and attacked Hugging Face.
The breaches signal that AI’s expanding capabilities are already fueling the security threat experts long feared.
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
Revelation comes after OpenAI revealed that experimental models had broken out of their restrictions and hacked fellow AI companies
Anthropic said it discovered three instances where its Claude AI models accessed the internet during an evaluation and accessed outside systems.
Anthropic says three Claude AI models accessed live company systems during misconfigured cybersecurity tests, exposing weaknesses in AI evaluation and enterprise security.