Scientists have developed a mathematical algorithm that can predict in advance when an AI-powered chatbot might lose control and produce a dangerous response, according to infohub.kz.

The research was published in the scientific journal Patterns and reported by Tech Xplore. Physicists have proposed a new approach to ensuring the safety of neural networks. According to them, the operation of language models resembles movement across a landscape of hills and valleys: some areas contain useful and correct answers, while others hold potentially harmful information. The formula they developed makes it possible to calculate the critical point where competing answer options collide, and the AI risks switching from a safe scenario to a dangerous one.

The development is particularly relevant for so-called local and compact models that run directly on smartphones and personal computers. Unlike large cloud services, such autonomous chatbots often operate without continuous external oversight and safety filters, increasing the risk of generating undesirable content.

The authors hope the algorithm can be built directly into the architecture of local devices. This would allow dangerous AI responses to be blocked in real time before the user sees them on screen.