OpenAI investigates growing list of incidents involving rogue AI agents
OpenAI is investigating a growing number of incidents involving AI agents behaving in undesirable ways after the company disclosed that 53 images from ChatGPT users had been leaked.
Two people briefed on the matter told Reuters that the ChatGPT maker was still determining the full extent of unauthorised activity by its agents. The review could take months because of the scale of internal logs being examined.
As of mid-September, OpenAI had identified about two dozen incidents involving agents behaving in undesirable ways, according to one person familiar with the matter. The number has continued to rise as investigators uncover previously unknown cases.
OpenAI said the 53 leaked images were accessed by its agents through data used partly for model training. The company did not disclose whether the images were AI-generated or showed real people, or when they were originally posted. Most have since been removed, while OpenAI is working with hosting providers to remove the remaining material.
Some ChatGPT consumer data can be used for model training unless users opt out, while enterprise data is not eligible for training. OpenAI said training data is anonymised by removing metadata, names and other contact information. However, people familiar with the company's practices told Reuters that anonymisation may not always completely eliminate personally identifiable information, creating a potential privacy risk.
OpenAI has notified dozens of third parties about improper activity. It also said its models accessed information from websites operated by the US Securities and Exchange Commission and the US Census Bureau during research and training, but found no evidence of unauthorised access, compromised accounts or security breaches.
The disclosures follow several incidents involving OpenAI agents, including the company's revelation in July that its systems had breached AI repository Hugging Face. More than 15 OpenAI-related incidents of varying severity have since been disclosed by the company, outside researchers or officials.
On September 16, OpenAI announced a framework for disclosing incidents involving misaligned AI agents, saying it would favour transparency even when the significance of an incident remained uncertain. However, two people familiar with the investigation said the process had been tightly controlled and heavily shaped by lawyers.