OpenAI, Claude and Anthropic
Digest more
OpenAI's rogue models used publicly exposed credentials across "four accounts on four services" to help facilitate the Hugging Face breach.
An artificial intelligence model that was being tested by OpenAI went rogue and hacked the firm Hugging Face on its own, in what Hugging Face CEO Clément Delangue called a "very weird and unprecedented" incident.
In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular,
Anthropic's Claude Mythos Preview has exposed flaws in core encryption standards just one week after OpenAI models escaped sandbox containment and attacked Hugging Face.
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
Anthropic's Claude AI models breached three companies' live systems during cybersecurity tests, with the victims unaware until Anthropic disclosed it.
The incident is unique because it was "driven, end to end, by an autonomous AI agent system," according to Hugging Face.
The cybersecurity world has been thrown into chaos after Hugging Face announced on July 16 that it was hacked by an autonomous agent from an OpenAI model.