OpenAI, Face
Digest more
OpenAI's rogue models used publicly exposed credentials across "four accounts on four services" to help facilitate the Hugging Face breach.
In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
Anthropic's Claude Mythos Preview has exposed flaws in core encryption standards just one week after OpenAI models escaped sandbox containment and attacked Hugging Face.
Revelation comes after OpenAI revealed that experimental models had broken out of their restrictions and hacked fellow AI companies
Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular,
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
The breaches signal that AI’s expanding capabilities are already fueling the security threat experts long feared.
Leaders should understand open and closed weight models, rather than defaulting to the most expensive option out of habit