HomeArticle

Anthropic $1.5 billion settlement: The case is closed, yet the legal issues are only just beginning

互联网法律评论2026-07-29 12:13
The $500 million settlement sets a new record, yet legal issues remain unresolved. Congressional legislation is the real variable.

On July 20, 2026, Judge Araceli Martínez-Olguín of the U.S. District Court for the Northern District of California signed a final approval order, under which Anthropic paid a $1.5 billion settlement for using pirated books to train Claude, covering approximately 482,000 to 500,000 works, equivalent to roughly $3,000 per work. This is the largest copyright settlement in the history of U.S. copyright — and also in the AI era — to date.

But the most significant implication of this ruling may lie in the issues it does not resolve: $1.5 billion is not the "price tag for fair use in AI training", neither Anthropic nor the copyright authors can be regarded as the winner of this game, and the legislation being advanced simultaneously by the U.S. Congress is the line that truly determines the future of AI training data.

I. The $1.5 billion price tag refers to "unauthorized acquisition from pirated sources" rather than "AI training"

To understand this settlement, we must go back to the 32-page summary judgment ruling issued by Judge Alsup on June 23, 2025. The judge, known for his "software engineer-like rigor", made a key dichotomy:

1. The training act itself: Inputting legally obtained books into the model to train Claude constitutes fair use; scanning and digitizing after purchasing physical books does not constitute additional infringement;

2. Downloading and permanently storing more than 7 million books from pirated shadow libraries constitutes "inherent, irredeemable infringement".

This settlement agreement does not overturn the previous ruling. On the contrary, the focus of the case now falls on Anthropic's acts of downloading and storing files from shadow libraries such as LibGen and PiLiMi.

In other words, $3,000 per work is the "clearing price for pirated data", not the "licensing fee for legal training". The Alsup ruling has made it clear — if the data source is legal, the training act itself may not need to pay copyright licensing fees at all.

Why did Anthropic prefer to spend 1.5 billion to settle? If the facts of the case are unfavorable to Anthropic, it will face far greater liability for damages — the maximum statutory compensation for willful infringement under U.S. copyright law is $150,000 per work, and the theoretical exposure of 7 million pirated books is close to $1 trillion. The $1.5 billion is an affordable number for Anthropic to buy off the sword hanging over its head.

II. "Non-binding precedent" — this is the greatest legal legacy of the case, yet it has been seriously underestimated

Anthropic and its supporters call this settlement a "victory" because the Alsup ruling stated that training AI constitutes fair use. But the legal reality is that due to the settlement rather than an appeal, Alsup's ruling will forever remain at the level of "district court pretrial opinion" and does not constitute any binding precedent.

Under the U.S. common law system of precedents, binding precedents come from final judgments of appellate trials, and pretrial orders at the summary judgment stage only have "persuasive force" but no "binding force"; and once a settlement is reached, the case never produces a final judgment on the merits of the dispute.

This means that Anthropic has won "a very weighty legal opinion", but not a "binding precedent". Moreover, the settlement agreement does not constitute a continuous license, and does not exempt it from future claims arising from model outputs or subsequent training methods, and 6 authors have chosen to opt out of the settlement and file separate lawsuits.

At present, dozens of AI copyright cases are being heard in U.S. federal courts. The Hachette v. Google Gemini case in the Southern District of New York, the Kadrey v. Meta case in the Northern District of California, the Thomson Reuters v. Ross case — each is being heard independently, and divergences have already emerged. Judge Chhabria explicitly reserved room for counterarguments of the "market dilution theory" in the Kadrey case, forming a substantive divergence within the same court from Alsup's radical position. These ongoing lawsuits by copyright holders against AI technology companies can still reach completely different conclusions in law.

III. $1.5 billion — It only looks like a huge sum

Looking past the intimidating total figure of 1.5 billion, the truth is: the group of authors is far from the winner of this settlement.

1. Only 2/3 of the $1.5 billion will actually go to the authors

This $1.5 billion will not only be used to pay copyright owners, but also cover other major costs that may drive the actual payment amount below the widely reported baseline of $3,000 per work.

Even after a significant reduction by the court, before the rightsholders receive any payment, approximately $122 million of this settlement fund has been allocated to legal fees, litigation expenses and administrative reserves. If calculated according to the original quotation of the legal team, this proportion will soar to a quarter to a third of the settlement fund — which means that before the authors see any money, a large chunk of the pie has already been cut away.

2. Is $3,000 per work reasonable?

This settlement agreement introduces a long-missing element in the debate over AI training data: market price. Based on the pricing of 500,000 works, the price of each work is approximately $3,000, which represents the first judicially recognized unit value assessment for the use of copyrighted books in AI training.

More than 91% of the authors and publishers bound by the settlement agreement have received compensation, but a total of 54 objections have been submitted to the court. The core attack points of these objections are highly consistent:

(1) The compensation standard is too low: Multiple objectors pointed out that $3,000 per work is "negligible" compared with the statutory maximum compensation of $150,000, and one opponent stated bluntly that members of the collective only received 2% of the statutory maximum;

(2) The distribution mechanism is unfair: In the objections summarized by Authors Alliance, the "default 50-50 split between authors and publishers" is explicitly criticized as unfair to small authors; for specific categories of books such as CC-licensed works, the position of the collective representatives is even contrary to the interests of the authors.

What is more thought-provoking is that 350 collective members have chosen to opt out of the settlement (involving 1,802 works), and they would rather give up the guaranteed bottom price of $3,000 per work to retain the right to sue Anthropic separately. Most of those who chose to opt out are large rightsholders or publishers "who own multiple works and have mature legal representation" — they have the capital to bet on larger compensation. Ordinary individual authors, by contrast, are left in the settlement pool, accepting the reality of receiving far less than $3,000 per work.

IV. The Congressional Legislative Path: Unlocking the Black Box

Before the court approved the above settlement agreement, on June 30, 2026, the Subcommittee on Courts, Intellectual Property, Artificial Intelligence, and the Internet under the U.S. House of Representatives Judiciary Committee held the "A Mid-Life Crisis? IP and the Internet After 40" hearing. When Professor Viswanathan of Columbia Law School was asked for her opinion on the "fair use" ruling in the Anthropic case, she explicitly opposed the broad fair use standard of the Alsup style, and refuted the "fair use defense" of AI companies one by one: the transformative use is "unclear" in law; training does impair the actual or potential interests of copyright owners in the licensing market; "the high cost of licensing" is not a valid defense — complex licensing practices such as film production have long proved the feasibility of large-scale licensing.

But the solution she proposed is "a light hand of regulation": to incentivize the formation of a licensing market rather than impose mandatory intervention. This line of thinking precisely corresponds to the design philosophy of the U.S. Congress around AI training data copyright recently — the U.S. Congress is building infrastructure through "transparency" and "verification rights":

(1) CLEAR Act: The Copyright Labeling and Ethics in AI Reporting Act (S.3813, introduced on February 10, 2026, under consideration by the Senate Judiciary Committee), with only one core obligation — mandatory disclosure. It requires generative AI developers to submit to the Copyright Office a "detailed summary of every copyrighted work in the training dataset" 30 days before commercial release, and the Copyright Office will establish a public database based on this information.

(2) TRAIN Act: The AI Network Transparency and Responsibility Act (H.R.7209/S.2455, introduced by the House of Representatives in January 2026), with the core mechanism of administrative subpoena — copyright owners can apply to the clerk of the district court for a subpoena, requiring AI developers to disclose whether their works have been used for training.

In essence, the CLEAR Act and the TRAIN Act are "right to know acts", which solve the black box problem of AI training data rather than the infringement characterization of AI training. This means that the compensation issue for copyright owners is left to two paths to resolve:

(1) Judicial path: After obtaining evidence through CLEAR/TRAIN, copyright owners file infringement lawsuits on their own, and the court will determine whether it is fair use and whether compensation is required (the Bartz settlement is the epitome of this path)

(2) Market path: Transparency forces the formation of a voluntary licensing market — AI companies will actively negotiate licensing fees with copyright owners to avoid litigation risks

V. The Case Is Closed, But Legal Issues Are Just Beginning

Judge Alsup's 32-page ruling is the most ambitious fair use argument for copyright law in the AI era — but it has never crossed the threshold of "binding force". What Anthropic bought with $1.5 billion is not the final victory of the "fair use" argument, but the chance to prevent it from being overturned by a jury.

But with this $1.5 billion, the authors did not win either. After deducting legal fees, litigation expenses and administrative costs, and then splitting 50/50 with the publisher, an ordinary author actually gets no more than $1,500 per work — the residual value after the work is massively stolen, discounted for risk and divided layer by layer.

This precisely explains why congressional legislation is accelerating. The CLEAR Act and the TRAIN Act are doing a more fundamental thing: unlocking the black box. When every copyright owner can confirm that their work is used through the public database, and can access training records through administrative subpoenas, the "fair use" defense will have to be re-stated in every subsequent case. The "collective bundled discount" mechanism of the Bartz type will be disintegrated, replaced by the accumulated exposure of long-tail individual lawsuits.

For large enterprises that are building or using AI, the real compliance problem lies in: Can AI business withstand this level of transparent scrutiny?

$3,000 per work is not the answer, it is just the beginning.

This article is from the WeChat official account "Internet Law Review", author: Zhang Ying, published with authorization from 36Kr.