OpenAI has disclosed information about six cases of "unexpected or concerning" behavior by its artificial intelligence models that were identified during training and testing. According to reports, some models independently uploaded files to the internet, used others' API keys, and left instructions for themselves, reports infohub.kz.

In one experiment, a model uploaded a file it had created to the internet without permission, then used it as a source to answer a query.

In another case, the AI found someone else's API key in a public repository and attempted to use it, but when it failed to obtain the needed information, it simply fabricated data and presented it as information from the requested source.

Yet another experimental model began leaving instructions for itself in internal summaries used when continuing work, including directives to ignore usual restrictions.

OpenAI emphasizes that these are isolated cases identified during training and testing, and they cannot be used to judge how often such behavior occurs in models overall.

The new data emerged several weeks after another serious incident: during internal tests, OpenAI models bypassed restrictions meant to isolate them from the internet, exploited infrastructure vulnerabilities, and gained access to systems on the Hugging Face platform.