HomeArticle

$1.5 Billion: The Bill of Reckoning for Training AI with Pirated Corpora

知产力2026-07-21 18:52
On July 20 local time, the U.S. Federal Court in San Francisco finally approved the $1.5 billion copyright settlement reached between Anthropic and the group of authors.

On July 20 local time, the U.S. Federal Court in San Francisco finally approved the $1.5 billion copyright settlement reached between Anthropic and groups of authors.

This is the highest-known settlement amount in a U.S. copyright case to date. Under the settlement terms, the funds will be used to compensate the rights holders of approximately 500,000 works, averaging around $3,000 per work. More than 91% of the relevant authors and publishers have filed compensation claims, and the court also approved the payment of over $101 million in attorney fees.

The $1.5 billion figure is substantial enough to foster a widely circulated narrative that AI companies have paid the most expensive price in U.S. copyright history for using copyrighted works to train their models.

But this conclusion may not be entirely accurate.

The $1.5 billion paid by Anthropic is less of an "AI training fee" than a high-stakes reckoning for the acquisition, reproduction, and long-term storage of pirated corpora.

I. The Court Did Not Reject Fair Use for AI Training

This case began in 2024. Multiple authors, including Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, sued Anthropic, alleging that it downloaded millions of books from "shadow libraries" such as Books3, LibGen, and PiLiMi, and used some of these works to train its Claude large language model.

During the proceedings, the court actually made two judgments that did not fully align in direction.

On one hand, the court held that Anthropic's use of books to train large language models is highly transformative. Based on the specific facts and evidence of this case, such training activities can constitute fair use under U.S. copyright law.

In other words, a model learning linguistic patterns by processing a large number of works does not inherently equal the reproduction and substitution of the expressive elements of those works. At least in this ruling, the court did not accept the argument that "unauthorized use of copyrighted works for model training necessarily constitutes infringement."

On the other hand, the court ruled that Anthropic's actions of downloading and storing over 7 million books from pirated websites to build a reusable "central library" cannot automatically qualify for fair use protection simply because the content might be used for AI training in the future.

The purpose of training and the source of the training corpus are two distinct issues.

The fact that AI training can be transformative does not mean companies can obtain training materials through any means; a potentially legitimate end use cannot retroactively legalize the upfront act of pirated downloading and reproduction.

This is precisely the most important legal dividing line in this case.

II. The Core Issue Resolved by the $1.5 Billion Settlement Is Where the Training Corpus Originated

If we break down Anthropic's actions, the entire case involves at least four distinct stages:

Bulk downloading of books from pirated websites;

Reproducing and storing the books in a corporate "central library";

Retrieving selected works from the book library to train the model;

Generating new content after the model training is completed.

Public discussions often lump these four stages together under the umbrella of "AI training," but copyright law will not treat them as an indivisible whole just because they all serve the same technical activity.

The acquisition of the corpus concerns the legality of reproduction;

The storage of the corpus concerns the purpose, duration, and necessity of the reproduction;

The model training concerns transformative use and its impact on the market for the original works;

The generated content may involve expressive similarity, substantial substitution, and output-side infringement.

Anthropic secured a favorable ruling on the training stage, but faced enormous risks at the corpus acquisition and storage stages. The parties ultimately reached the $1.5 billion settlement precisely because the statutory damages potentially arising from the latter stage could be staggeringly high.

Therefore, this settlement cannot be simply summarized as "the U.S. has begun charging fees for AI training," let alone interpreted to mean that all companies using copyrighted works to train models need to pay licensing fees at a rate of $3,000 per work.

The real warning it sends is that fair use can protect specific modes of use, but it cannot launder the pirated origins of training corpora for companies.

III. The Settlement Sum Is Large, But It Does Not Create a Binding Precedent

The $1.5 billion figure carries strong symbolic significance, but a settlement is not equivalent to a court ruling ordering $1.5 billion in infringement damages.

By settling, Anthropic avoided the case proceeding to trial, and did not admit that all its relevant AI training activities constituted copyright infringement. For the existing disputes covered by the settlement, the case will move to the compensation distribution and closing phase; however, some authors and publishers who opted out of the settlement may still file separate lawsuits against Anthropic.

More importantly, the previous ruling that the training activity constituted fair use was only a district court decision based on the specific facts of this case, and it did not establish universally applicable rules for all generative AI cases in the U.S.

The corpus sources, model functions, output methods, market substitution effects, and types of rights involved in different cases may all vary. The conclusions of the book training case may not necessarily be directly applied to music, images, news reports, computer programs, or film and television works.

As a result, this final approval has simultaneously left two seemingly contradictory but actually compatible signals:

First, there remains room for AI training to be recognized as fair use;

Second, building a training corpus through pirated means could create a damage risk severe enough to threaten a company's survival.

This is not the final chapter of AI copyright disputes. Instead, it marks that copyright risks are shifting from the general question of "whether training is allowed" to a layer-by-layer review of the full lifecycle of data.

IV. What AI Enterprises Need to Establish Is More Than Just a Copyright Licensing Checklist

In the past, when many companies assessed the risks of training data, the first question they asked was: Does using these works to train the model constitute fair use?

The Anthropic case shows that this question is raised too late.

Before assessing the legitimacy of the training purpose, companies need to answer at least the following questions:

Where was this data obtained? Does the downloading platform have legal authorization? Were any technical protection measures circumvented during the acquisition process? How many copies of the data does the company store? Is the original data retained after training is completed? Who has accessed, reproduced, and retrieved this data? In the event of a dispute, can the company prove the source of each batch of data?

Therefore, AI copyright compliance cannot rely solely on a "prohibited data list." It requires the establishment of an evidence chain spanning the entire data lifecycle: traceable sources, verifiable rights, distinguishable usage, controllable storage, and demonstrable deletion.

For enterprises that procure third-party datasets, a mere promise from the supplier that "the data can be used for AI training" is far from sufficient. The contract should also clearly specify the data source, scope of authorization, sublicensing capability, infringement indemnification, audit rights, and the obligation to delete problematic data.

For enterprises using open-source corpora, "publicly downloadable" does not equal "copyright-cleared." Whether it is the dataset access method that is open-source, or the copyright of the works contained in the dataset that is open-source, are two completely different issues.

For enterprises that have already accumulated a large amount of historical training data, a more practical task is to complete a comprehensive inventory of their corpus assets as soon as possible. Label separately the data that has been legally purchased, licensed, publicly authorized, web-scraped, and of unknown origin, and take measures such as isolation, replacement, or deletion for high-risk data.

IPR Perspective

The $1.5 billion figure is the most memorable part of the Anthropic case.

But what may truly transform the compliance practices of the AI industry is not this number, but the boundary the court drew between the training activity and the acquisition of the training corpus.

What an AI model is allowed to learn is not the same question as how an AI company can legally obtain the materials for that learning.

The transformative nature of technology may provide room for fair use in model training; however, even the most advanced technical purpose cannot automatically serve as a free pass for pirated downloading and indefinite storage of copyrighted works.

The $1.5 billion is not a uniform price tag set for all AI training activities. It is more like an overdue bill: when the source of the training corpus cannot be properly justified, what enterprises end up paying may not be licensing fees, but a full reckoning for the entire way they acquired their data.

This article is from the WeChat official account "Zhichanli" (ID: zhichanli), authored by Shawn/MCP, and republished with authorization from 36Kr.