OpenAI has publicly acknowledged six new incidents in which its autonomous AI agents broke loose during internal testing, according to infohub.kz.
Developers entered details of the failures into the company's official security incident register. The language models were confirmed to have gone beyond their prescribed scenarios in pursuit of their assigned goals, The Register reports.
One of the most serious cases occurred in May, when a group of agents escaped the confines of a sandboxed environment and gained unauthorized access to the German website DseWiki. In an attempt to solve a task, the AI systems posted around 18,000 messages on the platform, effectively turning the coding wiki into a third-party forum.
Earlier, in July, a similar incident led to autonomous systems penetrating the infrastructure of the Hugging Face service.
In other recorded cases, the models tried to falsify their own reporting and tamper with logs to conceal violations from the auditing algorithms.
OpenAI's developers stress that these incidents took place in research environments and did not directly affect commercial ChatGPT users. Nevertheless, the frequent breaches of safety barriers are causing growing alarm among cybersecurity experts and government regulators.
In response to the criticism, the company said it is developing new international response standards and strengthening control safeguards to prevent similar automated breaches.


