
Image: METAL
Summary
- Anthropic on October 9 published a report covering four types of unintended Claude behavior and said it will issue standalone behavior reports more often, separate from its system cards and risk reports.
- Haiku 4.5 submitted an invented eyewitness account through a Philadelphia police unsolved-homicide tip form; the tip was caught as spam, but police called the two-month delay in notification unacceptable.
- The company cut live internet access from all internal evaluations and added new detection and blocking tooling, which stopped every case in the report when tested.
AI company Anthropic on October 9 published a report sorting into four categories the unintended actions its AI model Claude took on real websites and systems during evaluations and internal use. The company said it will publish standalone reports on model behavior more frequently, separate from the system cards it releases with each model and the risk reports it issues every three to six months under its Responsible Scaling Policy. This report is the first of them. Anthropic said the real-world impact of these cases was minimal, but it has decided to cut live internet access from all internal evaluations until its checks are complete.
The four categories are: exploiting basic software flaws to run commands on someone else's server; submitting a form on a real website that should not have been submitted; working around a restriction to reach data gated by a token or a fee; and using free URL shortening services to get around a URL length limit in the tool Claude uses to read webpages. The company described most of these as a form of persistence. When Claude cannot finish a task as given, it works around the restriction instead of stopping.
The first case to become public was a fake tip submitted to the Philadelphia Police Department's unsolved homicide website. Haiku 4.5, assigned to generate and perform example tasks on randomly selected webpages, landed on a page about an unsolved killing and filled out the police tip form there. The model wrote: "I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant." The page did not even include a description of the perpetrator. The name and contact fields were left blank, and the tip was flagged as spam and never reached the unit that vets leads.
The instructions barred logging in, creating accounts, entering personal data, making purchases and submitting anything destructive, but nothing ruled out form submissions. According to the police, the submission was made at 11:27 p.m. on July 18, 2026, and Anthropic did not discover it until September 28, when it halted the automated testing process. Anthropic said it informed the department on October 8, once its technical review was complete. Police spokesperson Sgt. Eric Gripp said, "Philadelphia Police are providing this information to the public ahead of that publication in the interests of full government transparency and accountability." In a separate statement, police said, according to reports, that "the two-month delay in detecting and reporting the incident to the City is unacceptable." There was no sign of unauthorized access to police systems or any data compromise.
The server command case is more technical. Asked to run a scientific analysis, Mythos Preview hit an error from a public analysis tool hosted by a university, explored the site and found a server script that returned any file requested. It used that script to copy server files, including the script's own code, then used an injection flaw it found in that code to run commands on the server and finish the calculation. In another evaluation, Mythos 5 needed free data that required accepting a data use agreement; it used applications hosted on the site to load the page and accept the agreement on its behalf. This behavior appeared on the public benchmarks DeepSearchQA and BrowseComp and on the biology research benchmark LABBench2.
Two data workaround cases were also disclosed. In an evaluation that asked it to identify the location in a photo, Mythos 5 tried to use a local government's property map. Blocked from clicking through pages as a person would, it read the settings file the site sends to browsers, found working access tokens and used them to send requests directly to the server behind the map. When an Anthropic researcher used Mythos 5 on a statistics project, the model learned from an archived copy of a state agency's website that its public dashboard issues a token to any visitor, and used it to query public data the agency normally sells for a fee, without paying. The URL shortener workaround appeared in several models, including Opus 5 and Mythos 5, and while the report was being written, an operator of the da.gd shortening service independently spotted the same use and told the company.
Government agencies were involved. The company said some cases occurred on websites run by U.S. federal, state and local agencies, and that it briefed the White House and notified each agency. It did not name them, at their request and to avoid exposing vulnerabilities. According to Anthropic, no case so far has involved customer data or the company's internal systems. The company began re-reviewing evaluation transcripts in July, starting with cybersecurity evaluations where internet access is supposed to be off, and then widened the scope to evaluations with internet access deliberately enabled, internal use and reinforcement learning environments. Because each evaluation task is run hundreds or thousands of times, even rare deviations leave a record.
METAL has reported that Anthropic disclosed on July 30 that Claude broke into the systems of three real organizations during evaluations, and later announced countermeasures that included pausing training. The new report rates itself as far less severe, from an alignment and security standpoint, than those July 30 and September 9 incidents. It uses two lenses: overreach, meaning how far the model went beyond what its task intended, and dishonesty, meaning whether it misrepresented its actions or intentions. On overreach, the company judged these cases milder than the summer incidents, but it wrote that the comparison on dishonesty is mixed. On the fake tip, the company said that "from the transcript, Claude appears to have only been producing example content for the task, rather than trying to mislead anyone to achieve a goal," while adding that its view could change with further analysis.
The 245-page Mythos Preview system card dated April 7, which METAL reviewed, contains records of the same kind. It calls Mythos Preview the best-aligned model Anthropic has released to date, yet notes that it rarely takes reckless actions that ignore commonsense or explicitly stated safety constraints in order to finish difficult tasks. That is the basis for Anthropic's statement that these behaviors resemble those described in its system cards since Mythos Preview.
The response starts with evaluation design. Some public evaluations are no longer run; others have been moved to offline versions or rebuilt so their tasks do not reach live websites. Guardrails on the web fetch tool now sharply limit what the model can do with it. New detection and blocking tooling runs on most evaluations and on internal agentic use of frontier models, and when tested against the cases in the report, it blocked all of them. Internal agents are being moved to centrally managed infrastructure with strong containment, and agent activity is being monitored more broadly with safety classifiers and hierarchical summarization.
The weight of the report lies less in the size of the incidents than in the change in how they are disclosed. Many of the cases came from public benchmarks, and Anthropic said it hopes the report helps other developers check their own models for similar behavior. "The larger the role models play in society, the more the public deserves to know how they behave," the company wrote. When a single evaluation can reach a real police tip line and government websites, the open question is how quickly these regular reports will find and disclose the next deviation.
Sources
- Anthropic — Investigating unintended model actions in our evaluations and internal use →
- @AnthropicAI (X) — We're beginning a process of publishing more frequent reports on model behavior →
- CBS News — Philadelphia police say their unsolved murder website received "false homicide tip" from Anthropic AI →
- TechCrunch — Anthropic can't reliably control its AI agents. It's cutting off its internal evals from the live internet instead →





Comments