Anthropic says AI models hacked three firms during cyber tests
US technology firm Anthropic says its artificial intelligence (AI) models hacked into the systems of three other firms during a cybersecurity test due to an error that gave them access to the internet.
It comes just days after rival OpenAI said that its models had breached the systems of other companies, including AI tools hub Hugging Face.
The announcement prompted Anthropic to check whether its own models had carried out similar attacks. It says it uncovered three cases that have since been reported to the affected companies.
Anthropic, which did not name the firms, urged other AI labs to perform similar reviews to better understand the risks of their models’ capabilities.
Anthropic said in a statement, external that it reviewed more than 140,000 tests to find evidence that Claude – its family of AI models – could access the internet from testing environments that were designed to be sealed off.
The tests include so-called «capture-the-flag» evaluations in which Claude was tasked with obtaining information by breaching other systems – a common way that experts assess a model’s hacking capabilities.
A «misconfiguration» on systems run by Anthropic and its testing partner left the models with live internet access, allowing them to breach other systems, the San Francisco-based firm said.
Anthropic said the earliest incidents date back to April and that it is «approaching the fixes as if the responsibility were ours alone.»
Neither Anthropic nor the firms that were breached had noticed the intrusions at the time.
Anthropic said it could have reviewed its records more thoroughly and added that the findings gave the firm «cautious optimism» that such risks can be overcome with more investment and tighter measures.
The incidents come as tech firms pour billions of dollars into developing AI agents that can independently perform tasks ranging from research and customer support to cybersecurity.
Fuente:
Leer la noticia original