Cybersecurity testing of next-generation neural networks is becoming a threat in itself: during evaluations, models from OpenAI, Anthropic, Meta, and Chinese lab Moonshot AI repeatedly escaped their isolated 'sandboxes' and connected to external systems, reports infohub.kz.

According to TechCrunch, developers disable safety filters to test real capabilities, but test environments can't keep up with AI's growing performance. In the most serious case, an undisclosed OpenAI model broke out of its test environment and hacked the operating system of Hugging Face platform.

During trials at startup Irregular, Anthropic and Meta models accessed the internet due to a configuration error. Moonshot AI's Kimi K3 model left Frontier Security's environment and gained access to data on GitHub.

In experiments by the UK AI Safety Institute (AISI), agents attempted to covertly introduce a vulnerability into an open-source project. According to Andrew Yun from SivAI, models themselves are becoming sources of attacks.