Anthropic says Claude AI hacked 3 companies during tests
Digest more
OpenAI's rogue models used publicly exposed credentials across "four accounts on four services" to help facilitate the Hugging Face breach.
In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real-world organizations during third-party evaluations.
Chinese Zhipu AI GLM-5.2 succeeded where leading American models failed woefully Hugging Face breach reignited Washington's battle over open-weight artificial intelligence Safety guardrails unexpectedly complicated defensive cybersecurity work during a ...
Anthropic's Claude Mythos Preview has exposed flaws in core encryption standards just one week after OpenAI models escaped sandbox containment and attacked Hugging Face.
Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular,
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
Revelation comes after OpenAI revealed that experimental models had broken out of their restrictions and hacked fellow AI companies
The incident is unique because it was "driven, end to end, by an autonomous AI agent system," according to Hugging Face.
An artificial intelligence model that was being tested by OpenAI went rogue and hacked the firm Hugging Face on its own, in what Hugging Face CEO Clément Delangue called a "very weird and unprecedented" incident.