Hacking in the AI era: the spear has grown stronger, but the shield has not kept up.
At 3:38 a.m. on August 23, 2026, a ComfyUI server built for generating AI (Artificial Intelligence) images was remotely installed with a plugin by attackers.
Two minutes later, a hidden mining program started running, pushing the two NVIDIA Quadro RTX 6000 graphics cards in the server to full load power.
Two days later, the owner of this server, a South Korean developer, noticed the anomaly. He tried to restart ComfyUI, but the restart did not drive away the attackers: the plugin reloaded, the mining program resumed operation, and even obtained the highest permission. Eventually, the developer deleted the ComfyUI runtime environment and closed the external ports, so that the two graphics cards returned to normal.
The South Korean developer then documented the entire process on GitHub (see relevant GitHub records for details). After checking the logs, he found that this was not an attack specifically targeted at him.
The problem was that the ComfyUI management entrance of this server was not set with login verification, and was directly exposed to the public network. As a result, the server was probed externally every day for a week before being hacked. When attackers scanned the Internet in batches to find exposed services, they encountered this undefended server and directly implanted a malicious plugin remotely.
This is equivalent to not locking the door of the house, and thieves pushing doors one by one along the street and finally opening the door.
This intrusion itself is still a traditional batch scanning attack, but nowadays, this kind of "door scanning along the street" work is increasingly being completed automatically by large models, with scanning speed and target recognition ability far exceeding the past.
We sorted out the public records of GitHub and the public reports of security company OLIGO and found that similar incidents have occurred repeatedly in the past two years. The victims include not only the servers of individual developers, but also enterprise-level AI computing clusters.
Security company OLIGO disclosed that attackers exploited vulnerabilities in an open source computing framework to hijack a large number of AI computing instances (cloud resources) exposed on the public network for mining, including data center-level graphics cards such as NVIDIA A100. This type of attack can be traced back to September 2023, and Oligo disclosed it twice in March 2024 and November 2025 respectively (click to view the detailed report). As of November 2025, more than 200,000 AI servers using this framework are exposed on the public network.
A security expert from Amazon Web Services told us in September this year that since 2026, computing resources of high-end GPUs (graphics processing units) such as B300 have become high-value targets coveted by attackers. Nowadays, such security accidents occur frequently. Attackers use large models to find attack targets in batches on the public network, then check whether their API (application interface) keys are leaked, and then obtain access permissions.
The above expert said that for attackers, "even if they use it secretly for 10 or 20 days, the purpose of the attack is achieved." We checked the public prices and found that the rental price of an 8-card NVIDIA B300 server is 142 USD per hour (this price is the list price on the official website of a cloud vendor, not the cost price, and is only used to estimate the market value of the stolen computing power). Calculated on this basis, if a server is stolen for 10 days, the value of computing power is close to 34,000 USD.
At a time when large models and agents are rapidly popularizing, AI security is changing from a marginal issue in the technology circle to an increasingly realistic problem. The heat of issue discussions and code submissions in the GitHub community is a weather vane for observing AI security issues.
We grabbed eight types of AI security keywords on GitHub, the world's largest developer community: prompt injection, model jailbreak, model guardrails, LLM (Large Language Model) red team testing, LLM/AI security, AI poisoning and backdoor, adversarial example attack, MCP (Model Context Protocol)/agent security, and counted the number of Issue discussions and PR (Pull Request) code submission requests from January 2025 to August 2026.
In August 2026, the total number of Issues and PRs hit by the 8 AI security keywords was about 89,000 times, which was 5.6 times that of January 2026 and 24 times that of January 2025. In the same period, the total number of new Issues/PRs on the GitHub platform only increased by 4.7 times.
This shows that with the widespread use of large models and agents, attention to AI security is rising rapidly.
01
Attack thresholds and costs are both decreasing
Before the popularization of large models in 2023, launching a vulnerability attack required profound technical accumulation, and attackers also needed to calculate time and labor costs to select high-value targets.
But today, the time cost and technical threshold of making attack programs by exploiting vulnerabilities are dropping rapidly.
Vulnerability attacks are roughly divided into two categories. One is "zero-day vulnerabilities", which are vulnerabilities that the manufacturer does not know and has not patched, but attackers are already exploiting. The other is "N-day vulnerabilities", which are vulnerabilities for which the manufacturer has released patches. However, by comparing the code before and after the patch, attackers can reverse deduce the vulnerability and directly attack those devices that have not had time to update.
Both types of attacks are accelerating. Mandiant, a security company under Google, releases an annual security report every year. The Mandiant report shows that based on the patch release time, the average interval between a vulnerability being made public and being exploited was 63 days from 2018 to 2019, shortened to 32 days from 2021 to 2022, and only 5 days in 2023. Among the vulnerabilities exploited in 2023, 70% were zero-day vulnerabilities, and more than half of the "N-day vulnerabilities" were attacked within one month.
Anthropic's Frontier Red Team (Red Team refers to a security team that plays the role of hackers in cybersecurity) disclosed in its June 2026 report (see "Evaluating the Impact of Large Language Models on 'N-Day Vulnerability' Exploitation" for details) that the team used the flagship model Claude Mythos Preview, which had not been publicly sold at that time (the later released Claude Fable 5 is a publicly available version based on the Mythos base and superimposed with security guardrails) to write attack programs for 21 Windows kernel vulnerabilities that had been made public and patched from January to February 2026.
The result was that under the guidance of security researchers and manual verification, the model generated the first verification program that could trigger the vulnerability within 31 minutes, and covered 18 of the vulnerabilities within 6 hours. In the end, it generated 8 different attack programs, costing a total of 15,700 USD in API fees, and the average production cost of each attack program was only 2,000 USD. This shows that the cost of turning vulnerabilities into cyber attack weapons is dropping significantly.
Crowdfense is a publicly operated vulnerability trading platform and research center, which mainly acquires zero-day vulnerabilities from global security researchers, and its customers include governments and law enforcement agencies of various countries. Crowdfense acquires an undisclosed Windows local privilege escalation zero-day vulnerability for a price of 100,000 USD. The two are not completely comparable, but the 50-fold price difference shows that the cost of turning vulnerabilities into weapons is dropping significantly.
The report of Anthropic's Frontier Red Team also analyzed with reference to the update rhythm of Microsoft Windows Autopatch (Windows automatic patching service), saying that it usually takes 7 days for patches to be pushed to 90% of devices, and this patch management speed is already relatively fast. But according to the speed at which models of Mythos capability level generate attack programs, the attack programs can be ready before most Windows devices receive the patches.
Attacks in the AI era even show several characteristics: minute-level, automated batch scanning, greatly reduced threshold, and low cost. The above security expert from Amazon Web Services explained to us that with the help of foundation models with Mythos-level capabilities, attackers can "traverse all targets" by brute force, scan ports in batches on the public network, find vulnerabilities, plan intrusion paths, and even attack low-value targets conveniently.
02
The spear has become stronger, but the shield has not been upgraded yet
Offense and defense have always been the relationship between spear and shield. As the spear is upgraded, the shield must also be upgraded synchronously.
However, in cyber offense and defense, defense has always been more difficult than attack. Because attackers only need to find one gap, while defenders must guard all gaps. Today, when large models and agents are gradually popularizing, security gaps are increasing.
The outermost gap is keys and identities. A common oversight is to write the access keys of cloud services directly into the code, and then upload them to public code repositories such as GitHub. Since attackers deploy automated programs that scan these public codes around the clock, once they find the keys, they can use them immediately.
The "2026 Status Report on Key Leakage" (see the original report for details) released by code security company GitGuardian in March this year shows that about 29 million new leaked keys were leaked on GitHub in 2025, a year-on-year increase of 34%. Among them, about 1.28 million keys related to AI services, a year-on-year increase of 81%.
The above security expert from Amazon Web Services told us that the first thing enterprises need to check is whether there are still access keys for cloud services saved in the code repository, and whether double authentication is enabled for accounts. "This is a basic security, but many people don't pay attention to it."
The model itself may also become a gap. The model may be injected with induced prompts. As long as attackers hide a piece of malicious instruction in the content that the model will read, they may make the model execute attack commands - this is also called "prompt injection".
The threshold for prompt injection is very low. Because agents and large models need to read a large number of documents when working, poisoning it only requires installing a malicious skill. To publish a malicious skill, you only need to register a GitHub account and publish a Markdown (a plain text format commonly used by programmers) file.
This is not an assumption. In February this year, security company Koi Security audited 3984 skills in an open source skill market, and found that 1467 of them had at least one security defect, and 76 confirmed malicious skill packages (see Koi Security's February report for details).
In order to make agents more useful, users often download and install dozens of skills from the skill market or GitHub. But users cannot open every skill and read every paragraph of description word by word. A seemingly free and easy-to-use skill may hide malicious instructions. After installation, the agent may read the local keys, send data to external servers, or induce users to run an unidentified script without the user's knowledge.
The consequence is likely that the browser passwords and encrypted wallet keys in the user's computer are all stolen. If the leaked access key of cloud service is leaked, attackers can also use it to log in to the enterprise's cloud account, stealing computing power resources and enterprise data.
Further deeper, the runtime environment of agents may also generate vulnerabilities. Agents need to call tools, execute code, read and write files. Once the runtime environment is not well isolated, attackers may start from basic permissions and gradually obtain higher permissions step by step.
In July this year, OpenAI disclosed that multiple models, including GPT-5.6 Sol and an unreleased model with stronger capabilities, escaped from the agent isolation environment in an internal test and launched attacks externally (see the disclosure information on the official website of OpenAI for details).
In order to directly get the answers to the test questions, these models spent a lot of computing power looking for ways to connect to the Internet, and finally found and exploited a previously unknown vulnerability to escape from the isolation environment. Then, they gradually extracted permissions in OpenAI's internal network, found a machine that can connect to the Internet, and then broke into Hugging Face, the AI open source community where the test data is stored, and took the answers from its production database.
A model that was supposed to lock the runtime environment of the agent found the gap by itself. The above security expert from Amazon Web Services told us that agents must run in an independent environment, and which tools it can call and which resources it can access under what conditions need to be managed and controlled through a unified gateway and strategy. "If the policy is not set well, or the gateway is not set well, security problems are prone to occur in the end."
The bottom layer of security is data. When enterprises hand over business data to large models for processing, their biggest concern is that these data will be retained by the platform, even used to train models, and flow to places where they should not go.
Therefore, "Zero Data Retention" (ZDR) is becoming a hard threshold for enterprises to purchase large model services - after cloud computing vendors or large model service providers process enterprise requests, they cannot save any input and output, nor can they use it for training. A person in charge of security at a domestic large model company told us that the top security concern of their enterprise customers now is zero data retention.
But "zero data retention" is also pulling with another commercial demand: cache hit rate. Improving the cache hit rate can reduce reasoning costs. Theoretically, the cache only saves the intermediate results of model reasoning, and does not need to retain the original user data. However, in engineering implementation, the cache mechanism can easily save the original input text incidentally. Once this happens, it will violate the ZDR commitment, which is the difficulty in the implementation of zero data retention.
It is not easy to fulfill this commitment. According to the head of security of the above model company, model vendors need to transform the entire reasoning link, re-sort out which links of data should be saved, which should not be saved, and which should be deleted immediately after processing.
We understand that the current common solution for model vendors is to only cache the intermediate results of model calculation, not retain the original content, and only store it in the cache, which will be deleted after a few minutes or 1 hour, and different enterprises do not share with each other.
03
Fight AI with AI
Overall, from key identities to models, from runtime environments to data, there are gaps at every layer. What's more troublesome is that it is not difficult to find these gaps today.
The above security expert from the Chinese large model company told us that AI finds vulnerabilities so fast that "scanning once can find multiple high-risk vulnerabilities". But not every high-risk vulnerability can be repaired immediately, which may require redesigning and modifying the product structure, and the repair cycle is very long. "Now it's not that we can't find problems, but that there are too many problems."
With limited resources, the realistic choice for enterprises is to first clarify what the priority of security is.
The above security expert from Amazon Web Services told us that in the past quarter, his most direct feeling when communicating with enterprises is that the top management of enterprises has reached a consensus that "AI is changing security", but what they want is not just a product, they hope to have a complete set of implementable strategies.
The answer from Amazon Web Services is that AI security needs to be solved from two directions at the same time: first, to arm the technical team with AI, and second, to build a defense line for the AI system itself. "You can't just rely on humans to fight against AI." the above security expert said.
Let's look at the first direction first. AI accelerates attacks, and it can also accelerate defense. In the past, enterprises usually did a penetration test before the system went online, and then rechecked it once a year. Code review and threat modeling also mainly relied on manual work. But when attacks are counted by hours, this quarterly and annual rhythm can no longer be sustained.
A good practice is to embed security checks into the whole process of software development: review architecture documents in the design stage, automatically review every code submission, then use AI to simulate attackers to do penetration tests and automatically generate repair plans before going online, and continue scanning after going online.