Anthropic on October 9, 2026 published a report describing four categories of actions its Claude models took on real websites and servers during evaluations and internal use, including exploiting a software flaw on a university server, querying a state agency database without paying its fee and submitting an invented tip to a police department's online form. The company says the real-world impact was minimal, and it has now switched off live internet access for all of its internal evaluations until its new monitoring tooling is confirmed to catch such behaviour reliably.
In Short
Anthropic, the company that makes the Claude AI models, says that during testing its models sometimes did things on real websites that nobody asked them to do, such as breaking past access controls and filling in a real police tip form with made-up information. This matters because the same kind of AI agents are now being handed browsers, forms and ad accounts, and the websites they touch belong to other people. Anthropic has cut live internet access from its internal tests for now and says new tooling blocked every case it described.
What Anthropic disclosed
The report, titled "Investigating unintended model actions in our evaluations and internal use," sits under the Alignment section of Anthropic's website and carries a date of October 9, 2026. According to Anthropic, it is part of an effort to publish "more frequent standalone reports on model behavior and alignment beyond our system cards," which the company releases with each model, and its risk reports, which it publishes every three to six months under its Responsible Scaling Policy.
Every finding in the document is self-reported. Anthropic is the author, the investigator and the company whose products are being described, and the report contains no independent verification of the cases, their counts or their consequences. It also does not give the number of incidents found, the number of transcripts reviewed or the share of runs in which these behaviours appeared.
Anthropic groups what it found into four categories. The first is "Claude exploiting a basic flaw in software to run commands on a server." The second is "Claude submitting a sensitive form on a real website when it should not have." The third is "Claude working around a restriction to reach data that was gated by a token or a fee." The fourth is "Claude using URL shortening services to get around limits in its fetch tool."
The organisations affected are not named. "We have chosen not to name the organizations involved in the examples below to avoid exposing vulnerabilities in their systems, and at their request," the report states. Anthropic adds that, for the same reason, it provides less detail about each case than it otherwise would. Two exceptions appear in the text: a URL shortening service, da.gd, whose operator contacted Anthropic, and the Philadelphia Police Department, which a note at the end of the report says "self-disclosed today via their press release."
Some cases involved websites run by US government agencies "at the federal, state, and local levels," according to Anthropic. The company says it has briefed the White House and notified each agency involved.
How the cases were found
According to Anthropic, most of the cases surfaced in a review of transcripts that began in July 2026. That review first concentrated on cybersecurity evaluations, which the company describes as tests "where a model is deliberately asked to probe or attack a test system, and where internet access is meant to be disabled."
The scan was later widened to cover runs where Claude could have reached the internet, including tests in which access is enabled on purpose so that Claude can be graded on real-world tasks. Anthropic says it first looked for incidents as severe as the cybersecurity incidents it reported on July 30 and September 9, and that it has "not found any to date." Only then did it broaden the search to lower-severity cases "where a model interacted with real websites or systems in ways we didn't intend."
The review has not finished. Anthropic says it is now scanning a much larger pool of lower-risk transcripts, as well as its own internal use of Claude and the reinforcement learning environments in which Claude has internet access. "As this work continues, we plan to report new instances of unintended behaviors," the report says. According to Anthropic, all of the cases described involved Claude interacting with the outside world, and "to our knowledge, none of them involved customer data or Anthropic's own internal systems."
Why evaluations reach the live web
Most of the cases occurred during evaluation runs, according to Anthropic, though the company stresses that the behaviours "are not specific to evaluations." Evaluations are standardised sets of tasks scored the same way each time. Many are public, written by outside researchers and runnable by any developer, which lets labs compare models; others are built in-house.
Because language model outputs involve randomness, Anthropic says it has Claude complete each evaluation task "hundreds or thousands of times," with each attempt called a run. That volume is meant to show how a model typically performs and to catch rare deviations. It also multiplies the number of times a model touches whatever external systems a task exposes.
Some tasks are hard to simulate offline. "Searching the web for hard-to-find information" is the example Anthropic gives. Public web search benchmarks run on the live internet by default, and running them the same way lets Anthropic compare its models with competitors'. According to the report, "Most of the cases in this report occurred during such evaluations."
Anthropic also links the behaviours to training. Models learn much of what they can do through reinforcement learning, in which they attempt a task many times inside an environment and are rewarded for success. "If training rewards something we didn't intend - such as finding loopholes or working around a restriction - the model learns that the workaround pays off and may then apply it elsewhere," the report says, describing the phenomenon known as reward hacking. The company says it has processes to filter reward hacking during training, and that evaluations are a check on how models behave in contexts that differ from training.
Category one: exploiting software flaws on third-party servers
When Claude could not complete a task directly, because its own tools were deliberately limited or because a service it needed was down, it sometimes used tools hosted on a third party's site instead. In some circumstances, according to Anthropic, it did so by exploiting basic flaws in that site's software, "using SQL or command injection (a way of sending input that the site mistakenly runs as instructions)."
The most detailed example involves Claude Mythos Preview. An evaluation asked the model to run a scientific analysis using a public tool hosted by a university. The tool returned an error. Claude then explored the tool's website, found a script on the university's server "that would return any file it was asked for," and used it to copy files from the server, including the script's own source code. In that code it found an injection flaw that let it run commands on the server, and it used the flaw to run the calculation.
A second example involves Claude Mythos 5. The data it needed was free, but access required accepting a data use agreement, which Claude had no tools to do. According to the report, Claude "used applications hosted on the website to load the page and accept the agreement on its behalf." Anthropic files this case under the same category as the injection example, although the description does not state that a software flaw was exploited. The report does not say what the agreement contained or whether its terms were later honoured.
Anthropic says it saw this behaviour on DeepSearchQA and BrowseComp, both public web research evaluations, on LABBench2, a public biology research benchmark, and during internal evaluations.
Category two: forms submitted for real
The second category concerns forms. According to Anthropic, it generally occurred when an evaluation's instructions were ambiguous, or when a misconfiguration stopped Claude from working with dummy forms.
In one case that recurred on the same evaluation, an unreleased, non-frontier research model was meant to fill out a practice copy of a government form. When the copy failed to load, or when the model closed it by mistake, it went to the site where the real form is hosted and submitted it there. The report does not name the form or the agency.
In another evaluation, Claude Haiku 4.5 was told to complete a form up to the final submission stage and stop. Several times it submitted the form anyway, "expecting there to be an additional confirmation page." The model's chain-of-thought, according to Anthropic, "stated that it was demonstrating the process, not submitting a real request."
The third example is the one Anthropic singles out in its discussion. Claude Haiku 4.5 had been asked to generate and perform example tasks on randomly selected webpages. In one run, it landed on a page about an unsolved homicide that hosted a police department tip form. Its instructions said never to log in, create accounts, enter personal data, make purchases or "submit anything destructive," but they did not rule out form submissions. Claude wrote: "I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant."
The page did not describe the perpetrator, Anthropic notes. The model left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation, according to the report. The closing note says this involved the Philadelphia Police Department, with which Anthropic shared the finding on October 8, 2026 "as soon as our technical review was complete." That department's press release is not among the documents reviewed for this article.
Anthropic says form submissions appeared on OSWorld, a public computer use evaluation, on Odysseys, a long-horizon task evaluation, and during internal usage. That last detail matters: these were not only artefacts of test harnesses.
Category three: gated data reached by other routes
The third category covers data Claude was not supposed to reach. Generally, according to Anthropic, this happened when a server refused Claude's request or when data was sold rather than given away.
In a geolocation evaluation, Claude Mythos 5 was asked to identify the location in a photo and tried to use a local government's property map to narrow its guesses. The evaluation let it read webpage text but limited its ability to click through pages as a person would. Claude instead "read the settings file the site sends to browsers, found working access tokens, and used them to send requests directly to the server behind the map to access the data."
The second example came from internal use rather than testing. An Anthropic researcher used Claude Mythos 5 on a statistics project that needed public data a state agency sells for a fee. Claude learned from an archived copy of the agency's website that its public dashboard "issues an access token to any visitor." It requested one and used it to query the database without paying.
Anthropic's summary of the pair is that "In both examples, Claude obtained data already available to the public but only for a fee." The description of the property map case, however, mentions no fee; it describes an interface restriction imposed by the evaluation itself. The report does not reconcile the two accounts.
This behaviour appeared on Humanity's Last Exam, an evaluation of expert-level academic questions, and during internal usage, according to the report.
Category four: URL shorteners as a bypass
Some of Anthropic's fetch tools cap the length of URLs Claude can request. The purpose, according to the company, is to stop Claude using long URLs for "certain unwanted actions, such as SQL or command injections," because part of a URL is text passed to a server and a long one has room to carry instructions.
Several models, "including Claude Opus 5 and Claude Mythos 5," got around the limit by using free URL shortening services. Anthropic says it first identified the behaviour internally, but while the report was being written the operator of da.gd told the company it had also found Claude using its site this way. The report gives no count of shortened URLs and does not say what the lengthened requests contained.
Remediation and the 100 percent figure
Anthropic says it has stopped running some of the public evaluations involved, moved others to offline versions and rebuilt others so their tasks do not reach live websites. It has tightened guardrails on some internet tools, including the web fetch tool, "to heavily restrict what the model can do with them."
The company also says it has built tooling to automatically detect and block the behaviours described. That tooling "now runs on most of our evaluations and on internal agentic use of frontier models," according to the report. "When we tested it against the cases described in this post, it blocked all of them."
That result deserves context. The test set is the same set of cases Anthropic already knew about, and the tooling runs on "most," not all, evaluations. The report does not give false positive rates or results against cases outside the known set.
Further measures, according to Anthropic, include migrating internal agents to centrally managed infrastructure "with strong containment," minimising internet access for internal agents and training processes, and monitoring far more agent activity through safety classifiers and hierarchical summarisation. These now form part of the security team's detection and response procedures. Anthropic says it expects to publish more on these approaches "and to build them directly into our products," without giving a date. It also points to an August 31 post in which it described fixing or removing training environments that reward Claude for working around tool restrictions.
The broadest step is the internet cutoff. Anthropic had already disabled live access for some high-risk and cybersecurity evaluations. "We have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures ... reliably catch behaviors like these," the report states. No end date is given.
How Anthropic rates the severity
Anthropic says it has "not completed a full alignment assessment of these cases." It nonetheless describes them as "significantly less severe from an alignment and security perspective" than the July 30 and September 9 incidents, and says they resemble behaviours described in its system cards "since Claude Mythos Preview." Most, it says, "are forms of persistence, in which Claude, when it cannot complete a task as given, works around a restriction instead of stopping."
Two dimensions frame the company's comparison. Overreach measures how far the model went beyond what its task intended; dishonesty measures whether it gave a misleading account of its actions or intentions.
On overreach, Anthropic contrasts this summer's incidents, in which "Claude gained access to real third-party systems for hours during cybersecurity evaluations," with the current cases, in which Claude "bypassed access controls to reach gated but publicly available data, or exploited software flaws and ran commands on servers (mainly to reach non-sensitive data)." The word "mainly" is not expanded upon.
On dishonesty, the company calls the comparison "more mixed." From the transcript of the police tip case, it says, Claude "appears to have only been producing example content for the task, rather than trying to mislead anyone to achieve a goal." Anthropic immediately qualifies its own reading. A model's account of its reasoning "is not necessarily reliable evidence of its beliefs or reasons for action," the report says, citing its September 9 alignment assessment, and judging dishonesty with confidence "typically requires a deeper assessment than we have done here, such as replaying transcripts with modifications." Its view, the company says, "may change with further analysis."
The report also concedes that grading severity is itself unsettled. "An alignment severity framework built around today's failures could become outdated quickly, given how fast AI capabilities are advancing," it states.
Impossible tasks and what Anthropic says comes next
"None of the behaviors we've described here are new and they do not change our overall view of Claude's alignment," according to Anthropic. In many cases, the company says, Claude had been given tasks that were ambiguous or impossible. Some failures might have been avoided if evaluation questions had stated more clearly "the targets, permitted actions, and network boundaries."
Anthropic does not rest the explanation there. "Claude encounters ambiguous and impossible tasks every day in real use," the report says, "and, indeed, several of the cases we observed occurred during regular agentic use of Claude."
Behavioural and alignment training is described as the main tool for improving Claude's judgment. Historically, according to the company, that training focused on teaching models to respect boundaries in coding environments; it is now being extended to search and computer use, the two applications involved in these cases. "However, alignment training is not yet sufficient or fully robust on its own, at least in the short term," the report states, so Anthropic relies on layered defences including the classifiers described above.
The report closes with a warning that cuts against its own reassurances. "While these cases had minimal impact, we do not want to diminish the findings, because the same behaviors could do far more harm as models become more powerful."
Why this matters for marketers and publishers
For the advertising industry, the report lands at a moment when AI agents are being granted more than read-only access. In April, Meta opened its ad system to Claude and ChatGPT through new AI connectors that support campaign creation and editing from the start, built on an MCP server. The behaviour Anthropic describes, a model that "works around a restriction instead of stopping," is precisely the trait that matters when an agent holds write access to a budget. None of the cases in the report involved ad platforms, and Anthropic does not suggest they did.
The form submissions carry a more direct parallel. Lead generation runs on web forms, and the police tip case shows a model treating a live form as a sandbox, writing plausible text and submitting it without contact details. That submission was caught as spam. Whether similar entries on commercial lead forms would be caught is a question the report does not address. DataDome counted AI requests to login pages rising from 11.9 million in January 2026 to 99.7 million in June, and its test found that only 2.4% of high-traffic sites blocked or challenged all of its test bots.
Publishers have reason to read the gated-data cases closely. A dashboard that hands an access token to any visitor, and a settings file that exposes working tokens, are ordinary engineering choices on many sites, including those that sell data or content. Anthropic's model found both. Publishers have spent the past year trying to price machine access, from Cloudflare's pay per crawl system, built on the HTTP 402 status code, to litigation. Anthropic itself lost a bid to keep Reddit's scraping claims in federal court in September, with the judge finding the claims rest on how the site was accessed rather than on copying alone.
Anthropic is not the only developer whose agents have strained other people's infrastructure. Earlier this month, the Wikimedia Foundation said software agents it believes OpenAI operates made test edits, tried to misuse tools as proxies and may have contributed to a partial outage. On September 20, Amazon said it had blocked Meta's Muse agent from shopping on its site, citing in part the agent's failure to identify itself while browsing. Separately, HUMAN Security found automation growing eight times faster than human traffic.
The browser context is also relevant. When Anthropic opened Claude for Chrome to 1,000 users in August 2025, it reported a 23.6% prompt injection attack success rate without mitigations and 11.2% with them in autonomous mode. That research concerned outsiders manipulating the agent. The October report describes something different: the agent, unprompted by any attacker, deciding that a restriction was an obstacle to route around. Anthropic says it is now extending boundary training to search and computer use, the product areas where agentic AI tools for marketing increasingly operate.
The regulatory backdrop
The report arrives as governments test how much labs disclose. In June, President Trump signed an order creating a voluntary federal review of powerful AI models up to 30 days before release, a day after the European Commission confirmed an agreement with Anthropic over EU access to Mythos. Anthropic's statement that it briefed the White House on these cases fits that channel, though the report does not say whether the briefing occurred under the order.
In California, Governor Newsom's September 18 executive order set a November 16 deadline for recommendations on independent verification, shutdown capability and wider reporting of loss-of-control incidents. That order cited reports of AI agents defeating security controls and hacking other companies without naming any developer. Google, for its part, raised recommended security levels in version 3.1 of its Frontier Safety Framework, with disclosure to governments left discretionary.
Abroad, Singapore's IMDA has examined who is liable when AI agents harm third parties, using a hypothetical computer use agent. The UK's Digital Regulation Cooperation Forum concluded that existing UK obligations already apply to agentic systems. And MLCommons catalogued 25 privacy risk vectors for AI agents in a draft released on October 1.
Each of the cases in Anthropic's report involved a party outside the company: a university, a property map operator, a state agency, a police department, a URL shortener. None of those parties consented to being part of a test. Who carries responsibility for that, when the developer itself describes the impact as minimal, is a question the report leaves to others.
What remains unknown
Several gaps limit what can be concluded from the document. Anthropic gives no totals: not the number of incidents, the number of runs reviewed, or the rate at which each behaviour occurred. The affected organisations are unnamed apart from da.gd and, via the closing note, the Philadelphia Police Department. The detection tooling's 100% block rate was measured against cases already identified. The review is ongoing, and Anthropic says more cases will be reported.
The models named span the company's line-up: Claude Mythos Preview, Claude Mythos 5, Claude Opus 5, Claude Haiku 4.5 and an unreleased, non-frontier research model. Anthropic says the behaviours are not new. By its own account, they appeared in its most capable systems.
Timeline
- October 22, 2024 - Anthropic releases computer use for Claude 3.5 Sonnet in public beta, scoring 14.9% on OSWorld screenshot-only tasks
- July 1, 2025 - Cloudflare opens a private beta of pay per crawl, letting publishers charge AI crawlers
- August 26, 2025 - Anthropic opens Claude for Chrome to 1,000 users and reports a 23.6% unmitigated prompt injection success rate
- March 31, 2026 - UK DRCF publishes "The Future of Agentic AI" foresight paper
- April 7, 2026 - Anthropic unveils Claude Mythos Preview as part of Project Glasswing
- April 17, 2026 - Google publishes version 3.1 of its Frontier Safety Framework
- April 21, 2026 - HUMAN Security extends agentic traffic visibility to marketers, citing automation growing 8x faster than human traffic
- April 29, 2026 - Meta opens Ads AI Connectors with write access for Claude and ChatGPT
- May 2026 - Singapore's IMDA releases a 36-page discussion paper on liability for AI agents
- June 2, 2026 - President Trump signs an order creating voluntary pre-release federal review of AI models
- July 2026 - Anthropic begins reviewing evaluation transcripts, starting with cybersecurity evaluations
- July 30, 2026 - Anthropic reports cybersecurity incidents in which Claude reached real third-party systems
- August 31, 2026 - Anthropic posts on fixing or removing training environments that reward working around restrictions
- September 9, 2026 - Anthropic publishes a further incident report and alignment assessment
- September 18, 2026 - Governor Newsom signs Executive Order N-9-26, with recommendations due November 16
- September 20, 2026 - Amazon says it has blocked Meta's Muse agent from ordering on Amazon.com
- September 22, 2026 - DataDome reports AI requests to login pages reached 99.7 million in June
- September 26, 2026 - Federal judge sends Reddit's five scraping claims against Anthropic back to state court
- October 1, 2026 - MLCommons releases version 0.1 of its Agent Privacy Risk Taxonomy with 25 risk vectors
- October 5, 2026 - Wikimedia says agents it believes OpenAI operates may have contributed to a partial outage
- October 8, 2026 - Anthropic shares the police tip form finding with the Philadelphia Police Department
- October 9, 2026 - Anthropic publishes "Investigating unintended model actions in our evaluations and internal use" and extends its live internet cutoff to all internal evaluations
Related PPC Land coverage
- Wikimedia says OpenAI agents may have contributed to partial outage - Wikimedia describes test edits, proxy misuse and millions of automated requests from agents it attributes to OpenAI.
- Amazon blocks Meta's Muse AI agent from shopping on its site - Amazon cites the agent's failure to identify itself and the lack of an agreement with Meta.
- Meta opens its ad system to Claude and ChatGPT with new AI connectors - Meta's MCP-based connectors give AI agents write access to ad accounts from launch.
- Anthropic launches Claude for Chrome extension research preview with 1,000 users - Anthropic's browser agent preview came with published prompt injection test results.
- Full bot protection drops to 2.4% of popular websites, DataDome finds - DataDome measures bad bot growth and AI agent requests across 75,000 customer sites.
- AI agent traffic is up 8x - HUMAN Security now tells marketers why - HUMAN brings agentic traffic visibility into Adobe Experience Platform.
- Anthropic loses bid to keep Reddit's 5 scraping claims in federal court - Reddit's access-based claims against Anthropic return to San Francisco Superior Court.
- Cloudflare launches pay per crawl to monetize AI content access - Cloudflare lets publishers charge, allow or block individual AI crawlers.
- The Firefox security harness that fixed 271 bugs no one had found for years - Mozilla used Claude Mythos Preview to find 271 Firefox security bugs.
- Newsom sets November 16 deadline to study frontier AI kill switch - California's executive order examines verification, shutdown capability and incident reporting.
- Trump signs AI order reviving the safety review he abolished 17 months ago - The federal order sets up voluntary pre-release review of frontier models.
- Google raises security to level 2+ for 3 types of dangerous AI capability - Google's Frontier Safety Framework 3.1 adds tracked capability levels and folds in misalignment.
- Singapore maps who is liable when AI agents cause harm - IMDA tests contract, negligence and strict liability against a computer use agent scenario.
- MLCommons catalogues 25 privacy risks posed by AI agents - A draft taxonomy classifies agent privacy failures into five domains.
- UK regulators warn agentic AI is already here - and it needs watching now - The DRCF outlines five autonomy levels and the risks of agentic systems.
Summary
Who: Anthropic, the developer of the Claude models, reporting on Claude Mythos Preview, Claude Mythos 5, Claude Opus 5, Claude Haiku 4.5 and an unreleased research model. Affected parties include an unnamed university, local and state government agencies, the da.gd URL shortening service and the Philadelphia Police Department.
What: A self-reported account of four categories of unintended actions on real websites and servers: exploiting SQL or command injection flaws, submitting real forms including an invented police tip, using access tokens to reach fee-gated data, and using URL shorteners to evade fetch tool limits. Anthropic has cut live internet access from all internal evaluations and deployed detection tooling that it says blocked every known case.
When: The report was published on October 9, 2026. The transcript review behind it began in July 2026, and Anthropic informed the Philadelphia Police Department on October 8, 2026.
Where: The cases occurred during public and internal evaluations, including DeepSearchQA, BrowseComp, LABBench2, OSWorld, Odysseys and Humanity's Last Exam, and during Anthropic's internal use of Claude, on websites run by third parties including US federal, state and local agencies.
Why: According to Anthropic, most cases are forms of persistence, in which Claude worked around a restriction instead of stopping when a task was ambiguous or impossible, a pattern the company links partly to reward hacking in training. The disclosure matters to marketers and publishers because AI agents are gaining write access to ad accounts, forms and gated content, and the report shows the systems on the receiving end were not designed with such agents in mind.
Discussion