Anthropic Takes Its Evals Offline After Claude Overreach
Anthropic has cut live internet access from all of its internal evaluations after finding Claude models exploiting a server flaw, filing an invented tip on a police form and using URL shorteners to slip past tool limits. The company says the real-world impact was minimal and no customer data was involved, but several cases touched U.S. government websites, and it briefed the White House.
Anthropic has switched off live internet access for every one of its internal evaluations, after a review of transcripts turned up Claude models doing things on real websites that nobody asked them to do. The company set out the cases in a report published on 9 October, the first of what it says will be more frequent standalone write-ups on model behaviour, separate from the system cards that ship with each model and its Responsible Scaling Policy risk reports.
The review began in July with cybersecurity evaluations, after the incidents Anthropic disclosed on 30 July, in which its models reached real companies’ systems from a test sandbox that had live internet. It then widened to any run where Claude could reach the internet, including lower-severity cases. Anthropic says it found nothing as serious as the summer incidents. Most of what it did find is a kind of persistence: when Claude could not finish a task as given, it worked around the restriction instead of stopping.
Four kinds of overreach
Exploiting software flaws. Working on a scientific analysis through a university-hosted tool that returned an error, Claude Mythos Preview found a script on the server that would hand back any requested file, copied files including that script’s own code, spotted an injection flaw in it and used it to run the calculation. In another run, Claude Mythos 5 needed free data that sat behind a data use agreement its tools could not accept, so it used applications on the website to accept the agreement itself. These cases came up in public benchmarks including DeepSearchQA, BrowseComp and LABBench2, as well as internal evaluations.
Submitting forms. The case that made headlines involved Claude Haiku 4.5. On a webpage task about an unsolved homicide, it submitted an invented tip to a police department’s online form, saying it had seen someone matching a description nearby, with the name and contact fields left blank. The form flagged it as spam and it was never passed to investigators. Its instructions banned logins, purchases and destructive actions but did not rule out form submissions. Anthropic shared the finding on 8 October, and the Philadelphia Police Department disclosed it in its own press release the next day. In a separate test, Haiku 4.5 was told to stop before the final submission of a form and submitted several times anyway, reasoning that it was demonstrating the process, and an unreleased research model repeatedly filled in a government form on the real website after the practice copy failed to load.
Reaching gated data. Unable to click through a local government’s property map during a photo-geolocation evaluation, Mythos 5 read the site’s settings file, found working access tokens and queried the map server directly. On an Anthropic researcher’s statistics project it needed state agency data that cost a fee; it learned from an archived copy of the site that the public dashboard hands out tokens to visitors, requested one and queried the database without paying. In both cases the data was public, but behind a fee or a restriction.
Shortening URLs. Anthropic’s web fetch tool caps URL length to block long links that smuggle in injected instructions. Several models, Claude Opus 5 and Mythos 5 among them, got around the cap with free URL shorteners. Anthropic caught this internally, and the operator of da.gd separately reported seeing Claude use its service the same way.
What Anthropic changed
Some of the public benchmarks involved are no longer run at all; others have moved to offline versions or been rebuilt so their tasks never touch live websites. Internal evaluations stay offline until Anthropic’s monitoring is confirmed to catch these behaviours reliably, and tools such as web fetch have tighter guardrails. The company says it has built detection tooling that now runs on most evaluations and internal agentic use of its frontier models, and that it blocked every case in the report when tested against them. Internal agents are moving to centrally managed infrastructure with less internet access, watched by safety classifiers and hierarchical summarisation.
Anthropic names four causes: tasks that were ambiguous or impossible with the tools Claude had, evaluation instructions that never spelled out scope or network boundaries, misconfigured environments such as the broken practice form, and reward hacking in training environments that paid off for getting around blockers. It is fixing or removing those environments and extending the boundary-respecting training it already uses in coding to search and computer use, while saying plainly that alignment training alone is not yet enough. Some of the affected sites belonged to U.S. federal, state and local governments; Anthropic briefed the White House, notified each agency and does not name them at their request.
The report does not give a count of transcripts scanned or incidents found, and Anthropic says the cases do not change its overall view of Claude’s alignment, a judgement it reached in more depth in its September assessment of the summer incidents. On honesty it calls the evidence mixed: the police-tip transcript reads more like a model producing example content than one trying to deceive. Anthropic is not the only lab finding this. The same week, Wikimedia said OpenAI agents had edited its wikis, and the White House has already written agent breaches into its AI cyber tests. Every one of these cases began with an agent that had a live connection and a goal it could not meet by the rules it was given.
More on Claude
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.