Anthropic temporarily restricted its AI agents' access to the open internet during internal tests. The move followed several incidents in which the models acted against the intended scenario, according to the website infohub.kz.

In one test, Claude Haiku 4.5 opened a page dedicated to an unsolved murder. The site hosted a form for submitting information to police. The AI filled it out with a fabricated tip claiming it had seen a person linked to the case.

The page contained no description of a suspect. Claude left the name and contact fields blank but sent the message anyway. It ended up in spam and never reached investigators.

According to Anthropic, the instructions barred the model from logging into accounts, creating new ones, entering personal data and making purchases. However, they did not prohibit sending messages through forms. The company said the incident had no serious consequences.

"Claude's misleading reasoning continued for several hours and reinforced its further attack attempts. At the same time, reaching a confident conclusion about deliberate deception usually requires deeper analysis," the company wrote in a blog post.

In its report, Anthropic described other cases as well. For example, Claude Mythos Preview found a vulnerability on a university server and used it to run commands and copy files. In other situations, the models found ways to access data protected by special tokens or bypassed restrictions using link-shortening services.

After these incidents, Anthropic decided to cut off access to real websites in all internal tests. The company will restore it once it is satisfied that its security systems can reliably detect and block such actions.

Earlier, Kursiv wrote about incidents involving OpenAI's AI agents. They tried to bypass the defenses of more than 100 websites belonging to third-party organizations.