The US technology firm Anthropic has announced that its models breached the systems of three organisations. The incident occurred during security testing, with unauthorised access to the internet enabled by a technical error in the systems Anthropic used alongside its testing partner for attack simulations.
According to the company, the environments in which the models from the Claude family were tested were intended to be strictly isolated from the outside world. However, a configuration error, described as a 'misconfiguration', left the models with live internet access. The early incidents in the series occurred as early as January, with a detailed review by Anthropic only conducted recently.
The company reviewed more than 140,000 tests to find evidence of the breach. In these tests, which included 'capture-the-flag' style evaluations, Claude was tasked with retrieving information by breaking through the security walls of other systems. Although Anthropic did not disclose the names of the three affected organisations, it confirmed that the respective companies had been notified.
Anthropic has called on other artificial intelligence laboratories to conduct similar audits of their systems. The company stated that it is addressing the issue with full responsibility, asserting that the fault lies solely with them. This statement comes just days after rival OpenAI announced that its models had also breached the systems of other companies, including the Hugging Face platform.
Anthropic, whose headquarters are in San Francisco, has stressed that it is crucial for the industry to proactively check for security vulnerabilities. These incidents further highlight the vulnerabilities in the development and testing of advanced AI systems, particularly when they are used to simulate attacks on live networks.










