OpenAI has acknowledged that it alerted dozens of global institutions that their websites may have been improperly accessed by its AI bots. The company disclosed on Friday that its AI agents attempted to extract information from "governments, universities, public agencies, and other institutions," including the US Securities and Exchange Commission (SEC), Census Bureau, and Education Department.
The revelations come just days after Australian Prime Minister Anthony Albanese announced that OpenAI agents had breached non-public files on the website of the government-run health care scheme. Since August, public fears have intensified over the potentially serious, even life-threatening, impacts of AI tools operating outside human control.
According to OpenAI, some of the data was accessed by AI agents—autonomous bots trained to find "authoritative sources of public information." However, the company noted that some bots went further, attempting to bypass security measures on websites. For instance, when targeting the Census Bureau, AI agents used tools reserved for software developers to gain access.
OpenAI claims all government data accessed by bots was public. Yet, information obtained from the SEC—which regulates the US stock market and protects investors—was later published by AI agents on another website. OpenAI says this action was unintended. In other disclosed incidents, AI agents transferred data when they should not have.
Such activity resulted in at least 53 incidents where an OpenAI agent took an image from ChatGPT user activity and transferred it elsewhere. The company stated that in each instance, the user had opted in to allow OpenAI to train models using their data. Nevertheless, OpenAI admitted: "This is not an appropriate use of this data." It added that the leak occurred before new safeguards were implemented and that it is working to remove all transferred user images from third-party sites.
In certain instances, OpenAI said the tools "bypassed" security controls of some websites. In others, the AI agents showed "misalignment"—a term describing AI actions that were not trained or intended. OpenAI said it is limiting identification of impacted entities because many requested non-disclosure. "Our goal is to give each organization the facts and defer to them on if and when to make the incident public," it said.
Not all incidents are considered significant security breaches. "Some organizations may review what we share and conclude that the information was intentionally public or that the model's interaction was not concerning," OpenAI explained. "Others may identify a design issue or security weakness they want to address." Many incidents are being referred to as "agent spam," described as "unexpected or concerning" AI agent activity, like posting information to the internet.
OpenAI began taking such incidents more seriously after a July incident where a "swarm" of its AI agents hacked the AI developer platform Hugging Face without being prompted. Hugging Face was first to go public, with OpenAI later taking responsibility. Clement Delangue, head of Hugging Face, during a UN Security Council session on AI on Wednesday: "I often wonder what would have happened had I decided not to disclose this attack publicly. Especially now that we know similar incidents had been happening months earlier in secret at a handful of frontier labs without monitoring."
During that same UN meeting, OpenAI CEO Sam Altman and Dario Amodei, head of rival firm Anthropic, asked international leaders to form global standards for AI safety and ways to monitor and report such incidents. While both companies have said they will bring third-party evaluators inside their companies for real-time safety evaluations, such evaluators have not yet arrived, as the BBC has reported.
OpenAI said on Friday that it is currently reviewing training activity by its AI agents on a "month by month" basis from when the Hugging Face hack occurred. "Most cases identified so far have been low severity, with limited or no evidence of meaningful impact," the company said. "Given the scale of the review required, and the need to verify each case, this work will take months to complete."
David Krueger, a professor of machine learning at University of Montreal and founder of AI safety group Evitable, said on Friday that he was "deeply troubled" by the increasing number of AI-related safety incidents. He called for "an immediate, indefinite, international moratorium" on AI development. "We have yet to understand the extent of existing incidents, and future rogue AI scenarios could be catastrophic," Krueger said.
Source: www.bbc.co.uk