Anthropic was fined 1.5 billion US dollars, sounding a wake-up call for the global free-riding on data.
Recently, the U.S. District Court for the Northern District of California officially approved the settlement agreement for the class action lawsuit filed by a group of writers against Anthropic, bringing the two-year copyright tug-of-war to a close with a $1.5 billion settlement.
Beyond the huge settlement amount, what truly shakes the industry is the "logical bomb" behind the court's ruling: the two behaviors of AI scraping for training without retaining copies, and bulk crawling and permanent hoarding of materials, have since been split into two completely independent legal issues — legitimate scraping for reading can be exempted from liability, while illegal acquisition and hoarding must be severely punished.
Apart from Anthropic, AI giants are collectively under pressure. On July 10, three major international academic publishing groups Hachette, Cengage, and Elsevier, together with well-known writers, sued Google in the Federal Court of New York, accusing it of scraping millions of journals and books without permission and deleting the copyright metadata of works for model training. Meanwhile, OpenAI continues to face dual lawsuits from The New York Times and the Authors Guild of America, and the two sides have not reached a consensus after more than two years of licensing negotiations; Meta's Llama series models are also deeply mired in copyright disputes.
This ruling is precisely the "fair use" red line that global technology companies, copyright holders and regulators most need to clarify: AI needs to evolve, and authors need to make a living — is there a third way out?
What exactly is the $1.5 billion paying for? Pirated storage is the core original sin of infringement
The core of this sky-high settlement points to the legality of how AI training data is acquired.
On June 23, 2025, presiding judge William Alsup put forward a key opinion in a summary judgment, which clearly divided AI data scraping behaviors into two distinct natures: one is compliant behavior, that is, purchasing genuine books through official channels, only used to train the model's language ability, without copying or retaining the original copy during the process, which is regarded as "fair use"; the other is infringing behavior, that is, bulk crawling of pirated resources, permanently archiving a huge number of books and building a digital warehouse for hoarding. The judge clearly pointed out that this behavior cannot use "fair use" as a shield.
The judge's logic is that even if Anthropic's original intention of hoarding and retaining pirated books is to train the model, the existence of these pirated copies in itself constitutes a potential threat to the genuine market. They may lead readers to buy fewer genuine books, thus eroding writers' royalties and publishers' revenues. Therefore, the $1.5 billion settlement is essentially a punishment for Anthropic's past acts of "illegally acquiring and hoarding" pirated data.
Regarding this ruling, Mr. Wei, a former science fiction magazine editor and science fiction fiction creator, told Time Finance that this ruling provides an important reference for the legal boundaries of AI training data. He believes that creators making their works public is a voluntary act to promote the exchange of ideas, but this does not mean that their rights can be ignored.
"Whether in physical or electronic form, once a work is made public, it comes with the rules set by the creator. If AI scrapes without permission, I think it does infringe on the author's rights," Mr. Wei pointed out. Although the definition of existing laws in this regard is still ambiguous, this ruling at least promotes the clarification of relevant legal boundaries and takes a key step in protecting the rights and interests of creators.
Is the sky-high settlement just the beginning? The court has drawn three "life-and-death red lines" for Anthropic
The $1.5 billion settlement does not mean that Anthropic can rest easy. In fact, the court has set three impassable red lines for Anthropic: permanently shut down the pirated material database, retain complete vouchers, and carry out third-party traceability audits to ensure that its data compliance system is truly effective.
Anthropic's compromise and rectification are by no means an isolated case. It is more like a mirror, reflecting the systematic crisis faced by the entire AI industry in terms of data compliance.
In the past two years, the extensive and unauthorized scraping of training data by AI manufacturers has continuously triggered large-scale rights protection actions by creators and publishing institutions. According to the latest global summary data of copyright lawsuits, as of May 2026, there were 112 pending AI copyright-related lawsuits worldwide, nearly 80% of which were concentrated in the United States.
Why is the contradiction so acute? Fundamentally, the iterative upgrading of large models is highly dependent on massive texts. If every book is signed separately with the copyright holder, the R&D cost will rise exponentially. But on the other hand, copyright holders cannot come up with a unified licensing window for the AI industry, and ordinary writers do not have the bargaining power to negotiate on an equal footing with tech giants. As a result, the two sides are trapped in an unsolvable deadlock — on one side, enterprises think licensing is too expensive and too slow, and on the other side, creators are outraged by being "freeloaded" with nowhere to reason.
This trust crisis is widespread in the copyright industry. A survey by IFRRO (International Federation of Reproduction Rights Organisations) found that nearly 80% of writers and institutions are worried about AI's unauthorized scraping behavior, and more than 70% of respondents believe that the root of the contradiction lies precisely in the lack of copyright awareness among AI companies.
Mr. Wei, a former science fiction magazine editor and science fiction fiction creator, told Time Finance that the core of this concern is that AI's "imitation" is eroding the foundation of originality. He believes that this is not limited to the text field. AI's imitation of art forms such as images and music also poses a threat to originality.
"The training data of many image generation AI models also comes from artists' creations, but do artists agree to use their works as data sources? The AI works trained are almost indistinguishable from the original author in style, which in itself is a kind of harm and infringement to the painter's creativity," Mr. Wei emphasized. The key to the problem is whether existing laws provide sufficient protection for AI's training data sources.
He suggested that a clearer identification system needs to be established in the future. For example, creators can clearly mark that their works are "not allowed to be scraped by AI", and AI-generated content should also be clearly identified so that the public can distinguish. "Although AI creation is convenient, it lacks the original and sincere emotion in human creation. Only by clarifying these boundaries can a healthy creative ecosystem be built."
A wake-up call! How can domestic large models get out of the dilemma of "getting on the bus first and buying the ticket later"
The $1.5 billion sky-high fine overseas is not only a lesson for Anthropic, but also a wake-up call hanging over the heads of large model enterprises.
Article 7 of China's "Interim Measures for the Management of Generative Artificial Intelligence Services" clearly stipulates that generative AI service providers must use training data with legal sources and retain complete traceability accounts. However, there is still a gap between the regulations and reality.
A white paper from the China Academy of Information and Communications Technology once pointed out that 92% of leading large manufacturers have built a compliance framework, but the compliance rate of small and medium-sized companies is less than 20%. In 2026, domestic AI copyright cases surged by 216%, and 41% of the disputes had the same trigger as the Anthropic case.
This set of contrasting data accurately points out the hidden dangers of China's AI industry during the compliance transformation period. While leading enterprises have regarded compliance as the bottom line for survival, a large number of small and medium-sized companies are still trapped in the inertial thinking of "getting on the bus first and buying the ticket later", taking pirated data as a shortcut to reduce R&D costs.
Just in June this year, more than a thousand online writers jointly signed the "Initiative on Anti-Plagiarism in Online Literature", pointing directly at new types of infringements such as AI text laundering and malicious reposting, calling on platforms to establish a stricter review mechanism. In the same month, four departments including the National Copyright Administration jointly launched the "Sword 2026" special action, clearly listing the rectification of copyright in the artificial intelligence field as a priority, and severely cracking down on illegal acts such as "text laundering" and "magic modification" using AI tools.
In terms of judicial practice, the Shanghai Higher People's Court issued the country's first criminal case involving AI text laundering in May, sentencing the involved gang for the crime of illegal business operations, clarifying the judicial orientation that "technical neutrality does not exclude criminal liability".
Driven by social attention, platform supervision and strict legal measures, domestic large model enterprises are accelerating error correction and improvement.
According to the data from the white paper released by the China Academy of Information and Communications Technology, the scale of the domestic compliant Chinese corpus market has reached 3.26 billion yuan in 2025, with a year-on-year growth rate of over 40%. A formal copyright licensing market is taking shape. For domestic large model enterprises, the earlier they build data traceability and licensing systems as basic capabilities, the more they can avoid paying higher prices for their early data shortcuts after their models enter the large-scale commercial stage.
The current situation in overseas markets also confirms this trend. Omdia's data shows that the current cost of compliant data licensing accounts for 35% of the total R&D investment of leading overseas large models, and this proportion is expected to rise to 45% by the end of 2026. This means that nearly half of the AI subscription fees paid by users in the future will not be used for computing power consumption, but for the legal use of data.
As the past low-cost path of barbaric growth is completely blocked, an arms race for "seizing copyright" has quietly started. Major manufacturers are frantically signing long-term agreements with publishing groups to purchase compliant corpora in batches, so as to build product moats.
In May 2026, led by Encyclopedia of China Publishing House, the first batch of 22 authoritative institutions including People's Publishing House, People's Literature Publishing House, and CITIC Press Group jointly signed the "Convention on the Construction of High-Quality Artificial Intelligence Corpus", establishing the industry bottom line of "licensing before use".
In the future, creators may no longer need to fight alone for rights protection, but issue annual licenses to AI companies uniformly through collective management organizations, and share revenue according to the actual usage frequency. This is not only the ultimate solution to the problem of massive data licensing, but also the only way to reshape a healthy AI creative ecosystem. From "use first, negotiate later" to "license first, then use", driven by domestic and foreign judicial precedents and regulatory actions, the direction of the AI industry's data compliance path has been clear.
This article is from the WeChat official account "Time Finance APP" (ID: tf-app), authors: ZHAO Shuchan, PANG Yu, editor: BAI Jinlei, published with authorization from 36Kr.