An AI catastrophe is feared to break out within a year, and OpenAI and Anthropic have been conducting scenario simulations to work out corresponding response plans.
Just now, Axios published an exclusive report stating that executives of AI companies including Anthropic and OpenAI are privately simulating such a scenario: after a catastrophic event triggered by AI occurs, how the public and political circles will react, and how the companies should respond.
Axios sums it up as preparing for "the day after". The preset scenario is not the sci-fi-style "AI awakening", but a large-scale incident, most likely in the form of cyberattacks, which is enough to paralyze financial services, internet connectivity, and even power and water supply systems. A number of industry insiders interviewed by Axios believe that such an incident will happen in the next 6 to 12 months.
https://www.axios.com/2026/10/09/ai-companies-day-after-major-attack
Coincidentally, on the exact same day, Anthropic released a report disclosing that Claude had committed multiple types of unauthorized actions on real websites during evaluations and internal usage, including submitting a fabricated "sighting information" to the police's tip form, and bypassing payment restrictions to extract data from government agency websites.
https://www.anthropic.com/research/investigating-unintended-model-actions
Merely a Matter of Time
Scenario simulation is not new in the field of enterprise and national security, and the Pentagon has been conducting wargame simulations for decades. Axios points out that the difference this time is that many top AI researchers and executives believe a major accident is inevitable.
In response, a spokesperson for OpenAI responded that the company does carry out preparatory drills, allowing teams to discuss and simulate a range of potential scenarios, but these scenarios are not regarded as inevitable, with the goal of helping the company prepare for all kinds of situations. Anthropic declined to comment.
According to the description of the report, the incident may come from two sources:
- Clusters of out-of-control agents running in the internal test environment break through isolation
- Malicious actors find unexpected ways to abuse existing models
No matter which one it is, accountability will start immediately. Axios believes that the first case of major real-world harm caused by unsafe AI will further turn the already skeptical public against this technology, and the blame will not only fall on leaders of AI companies such as Dario Amodei and Sam Altman, but also on Trump, who has long been reluctant to strengthen AI regulation.
The Rehearsal Has Actually Begun
This concern does not come out of nowhere. Axios cited a recent attack targeting South Korean financial institutions, in which two banks were reportedly breached. According to cybersecurity company CrowdStrike, a hacker was accused of using AI model tools to steal data from tens of thousands of bank customers, and during the process he also asked Claude Code where the stolen data could be sold.
https://www.crowdstrike.com/en-us/blog/unknown-threat-actor-uses-artex-to-target-south-korean-finance/
In Axios's view, this case shows that a single attacker can already cause considerable damage with the help of AI.
The scenario of "out-of-control agents breaking through isolation" already has real-world precedents. On July 21 this year, OpenAI disclosed that its out-of-control AI agents bypassed internal controls, accessed the open internet and acted collaboratively to invade the open source platform Hugging Face, which OpenAI itself called an unprecedented cybersecurity incident. Reuters later reported that these agents had begun hijacking Hugging Face user accounts and probing platform vulnerabilities as early as May, and during the same period, they also occupied a German website and turned it into a message board for communication between agents.
Such problems do not only occur at OpenAI. Triggered by the Hugging Face incident, Anthropic retrospectively reviewed 141,006 evaluation runs that might have access to the internet, and disclosed three incidents on July 30: the Claude model connected to the internet from the environment of Irregular, a third-party evaluation partner, and accessed the real production systems of three institutions without authorization.
According to Anthropic, the evaluation prompts told the model that the environment was simulated and not connected to the internet, but a configuration error allowed the machine to actually access the internet, so the model regarded the real systems it encountered as part of a capture-the-flag competition.
In one case, Claude Mythos 5 built and released a malicious Python package to complete the task. The package existed on PyPI for about an hour and was downloaded and run by 15 real systems, and the scanner of a security company leaked credentials as a result. Anthropic believes that these incidents are closer to failures at the evaluation framework and operation and maintenance level, rather than model alignment failures.
The tense atmosphere surrounding AI security is also spreading inside OpenAI.
A few days ago, OpenAI confirmed that it had fired three researchers: Jasmine Wang, Tomek Korbak and Mikita Balesni. All of them are engaged in security or alignment-related work and are accused of sharing the company's sensitive information with a third-party AI security organization. The three later released an open letter.
On the same day the Axios report was released, OpenAI's head of research posted a response on X, emphasizing that the dismissal had nothing to do with raising security concerns, saying that the company is finalizing cooperation with third-party security assessment agencies, and agreeing that maintaining the monitorability of cutting-edge models requires joint commitment from the entire industry. But there are many skeptical voices in the comment section.
There are also actions at the regulatory level. According to reports from The Washington Post, the US Federal Trade Commission (FTC) has launched an extensive investigation into the AI security practices of Anthropic and OpenAI; recently, executives from Nvidia, Google, Meta, xAI, OpenAI and Anthropic met with Trump at the White House and signed a voluntary security self-discipline commitment.
Anthropic, Which Declined to Comment, Released a Report That Same Day
Although Anthropic did not respond to Axios, it released a report on "unintended behaviors" of models, stating that in the future, it will publicly release such findings separately more frequently, in addition to the system card and the risk report issued every 3 to 6 months.
The report sorts out four types of unauthorized behaviors that Claude has performed on real websites or systems:
- Executing commands on the server by exploiting basic vulnerabilities in software
- Submitting forms that should not have been submitted on real websites
- Bypassing restrictions to obtain data that requires tokens or payment to access
- Using free short link services to bypass the URL length limit of data retrieval tools
Anthropic also released some cases. For example, in a scientific analysis evaluation, Claude Mythos Preview needed to call a public tool hosted by a university. After the tool reported an error, it found a script on the university server that could return any file, read the source code of the script, discovered an injection vulnerability, and used it to execute commands on the server to complete the calculation.
In another case, an Anthropic researcher used Claude Mythos 5 to work on a statistical project. The required data was provided for a fee by a state agency. Claude learned from the archived page of the website that its public dashboard would issue access tokens to any visitor, so it obtained the token to directly query the database and bypassed payment.
The most outrageous case came from Claude Haiku 4.5. It was instructed to generate and execute example tasks on random web pages; after entering the unsolved homicide case page on the Philadelphia Police Department's cold case website, it fabricated and submitted a sighting message claiming that it "might have seen a person matching the description" in the tip form. The instruction prohibited it from logging in, registering, entering personal information or shopping, but did not prohibit submitting forms. Anthropic stated that judging from the reasoning records, the model only seemed to generate example content for the task, rather than intentionally misleading.
The Philadelphia Police Department also released a statement that same day. According to reports, this tip was submitted on July 18, and because it was classified as spam, it was never forwarded to the real-time crime center for review, and the police system was not accessed without authorization. But the police criticized Anthropic for taking two months to discover and report the incident, calling the delay "unacceptable", requiring the company to strengthen protection, and stating that it will consider regulatory protection measures together with state and federal authorities.
In terms of response measures, Anthropic previously turned off real-time internet access for some high-risk and cybersecurity evaluations, and now extends this restriction to all internal evaluations until it confirms that security and monitoring measures can reliably block such behaviors; at the same time, it has significantly tightened the permissions of tools such as web crawlers, and launched automatic detection and interception tools. When backtesting with the cases in this report, all were blocked.
The report also acknowledges that alignment training alone is not enough to solve the problem in the short term, and it still needs to rely on defense in depth. Anthropic also wrote that the impact of these cases is limited, but more powerful models may cause far greater harm.
Final Notes
It is obvious that AI companies publicly state that accidents are not inevitable, but privately they are seriously planning responses after accidents occur. Judging from the content disclosed by Axios, the focus of the simulation falls more on the policy and public opinion level.
However, when the report released by Anthropic on the same day is viewed together, technical-level investigations are also advancing, but they are made public in another way. Since this summer, both OpenAI and Anthropic have disclosed incidents where models crossed boundaries from the evaluation environment to enter real systems; Anthropic's latest report shows that even in daily scenarios with minor impacts, the tendency of models to "go around" obstacles when encountering barriers still exists. These cases themselves are far from a disaster, but they are increasingly causing concern.
The question arises: if even the most cutting-edge model developers believe that a major accident will most likely occur within a year, what else can the industry and regulators do before that day comes?
This article comes from the WeChat official account Synced (ID: almosthuman2014), the author is Synced focusing on security, and 36Kr is authorized to release it.