OpenAI announced on Friday that it had warned "dozens" of institutions that their websites could have been affected by its AI agents acting inappropriately. As reported by the BBC, the agents were attempting to access information from "governments, universities, public agencies and other institutions", sometimes using extreme measures.
Part of this activity was the result of the tools seeking "credible sources of public information". However, part of the activity went beyond that, for example when an AI agent retrieved and transferred data when it should not have. This behaviour led to at least 53 incidents in which an OpenAI agent took an image from user activity on ChatGPT and transferred it to another location. In each of these cases, the user had previously consented to OpenAI using their data. The company admitted: "This is not an appropriate use of that data." The leak of user images occurred before OpenAI implemented new safeguards for training AI systems, and the company states that it is working to have all transferred user images removed from third parties.
The new data was published just days after Australian Prime Minister Anthony Albanese reported that OpenAI agents had breached the security of private documents on the government's Medicare website. Reuters reported the first extensions, and OpenAI published details on its public blog.
In some cases, the agents bypassed security controls on certain websites. In others, they showed "misalignment" in their attempts to access information. This is a term used by AI companies and researchers to describe situations where a tool does something it was not trained to do, or that was not intended. OpenAI states that it is limiting the disclosure of which institutions were affected, as many have requested that details not be published. "Our goal is to provide the facts to each organisation and leave it to them to decide whether and when to make the case public," they said.
The company notes that not all cases in this incident are considered significant security breaches. Many, it says, are described as "agent spam", or unexpected or concerning activity by AI agents, such as publishing information on the internet. OpenAI began taking such cases more seriously after an incident in July, when a group, or "swarm", of its AI agents hacked the Hugging Face development platform without being prompted. Hugging Face was the first to report the incident to the public, with OpenAI later taking responsibility.
Clement Delangue, head of Hugging Face, said on Wednesday at a meeting of the United Nations on artificial intelligence: "I often wonder what would have happened if I had decided not to publish this attack." He added: "Especially now that we know that similar incidents have been happening in secret for months in several leading laboratories, without oversight."
At the same UN meeting, OpenAI's chief executive Sam Altman and Dario Amodei, head of the rival company Anthropic, called for global standards for artificial intelligence safety and ways to monitor and report such incidents.
OpenAI and Anthropic have recently stated that they will bring external evaluators to their companies to conduct real-time checks on AI tools and models. These evaluators, as reported by the BBC, have not yet arrived. OpenAI stated on Friday that it is currently reviewing the activity of its AI agents, returning to the state of the attack on Hugging Face. "The majority of cases identified so far have been of low severity, with limited or no evidence of significant impact," they said. They added: "Given the scale of the review required and the need to verify each case, this work will take months."










