HomeArticle

Sony and Warner are besieging Anthropic — How will the copyright war in the AI era be fought?

爱范儿2026-09-06 15:13
You got rich and didn't even take me along?

When talking about companies that are most fond of advocating "principles" in the AI sector and even the entire tech community, Anthropic is definitely the first name that comes to mind.

For the sake of its principles, the so-called "Company A" would even rather refuse to make money from you.

But probably because earning less money and paying out extra money feel completely different, when looking for training data back then, Anthropic unhesitatingly abandoned its principles and frantically imported tons of pirated materials.

Coincidentally, for music publishers that are used to waiting for content to mature before reaping profits, such behavior is exactly what they have been expecting.

On August 28, 35 music publishing entities composed of Sony Music, Warner Chappell and their affiliates filed a lawsuit against Anthropic, its CEO Dario Amodei, and co-founder Benjamin Mann in the U.S. District Court for the Northern District of California.

Considering that Universal Music had already sued Anthropic in the same court earlier this year, all three major global music groups have now joined this "general offensive".

What Exactly Did Claude "Take"?

As we all know, Anthropic is not Suno, and Claude cannot generate music either. So the focus of this dispute is not audio, but lyrics and music scores. This relates to some facts that have already been disclosed in the Bartz book copyright case last year —

In previous years, to train the Claude model, Amodei approved the team to download at least 7 million pirated books from pirated websites, and crawled a large amount of text on the Internet through web crawlers at the same time.

Photo | Bloomberg

Music companies believe that among these pirated books, there are at least hundreds of song collections, lyric books and music scores. Combined with the lyrics crawled from websites such as LyricFind through crawlers, the number of works whose copyrights have been infringed may reach tens of thousands.

Considering that a maximum of $150,000 in compensation can be claimed for each work that is intentionally infringed, the final amount of compensation claimed by Sony, Warner and other parties may reach as high as billions of dollars.

In fact, in the aforementioned Bartz book copyright class action, Anthropic has already paid a $1.5 billion settlement for its act of downloading pirated books.

Photo | Times: Andrea Bartz, one of the plaintiffs in the book infringement case

But the settlement is just the beginning. The willingness to settle sends a signal to the publishing industry that:

Anthropic can afford to pay, and it has the ability to pay.

As a result, numerous copyright holders involved in those 7 million pirated books are destined to come to claim compensation one after another, including several music giants that are the focus of today's event.

Training Is Allowed, But You Have to Pay for It

Contrary to the intuitive impression of many people, if Company A only properly acts as Kirby, "devouring" various materials and converting the text into the capabilities of Claude, this behavior itself is most likely not illegal.

From the previous precedent involving Google Books scanning tens of millions of books to build a directory index, to the court's judgment in the Bartz case on Anthropic's use of legally purchased books to train the model, U.S. courts have repeatedly accepted a view that:

Large-scale reproduction can sometimes be classified as "fair use".

Taking AI as an example, the model usually only learns the "knowledge" in the books, but almost never outputs the original text. What is presented to users is the result of re-narration after understanding, which usually cannot replace the value of the original book. In the legal sense, this belongs to the category of "highly transformative".

Photo | Electronic Frontier Foundation

In other words, styles and ideas themselves are not protected, and only specific expressions are the objects protected by copyright law.

From the perspective of the reproduction behavior itself, the model only splits the text and converts it into statistical associations encoded in forms such as vectors. It does not distribute new copies to the outside world, so it is hardly considered an infringement.

As for the act of scanning purchased physical books and then destroying them, at most we can criticize this practice as immoral, but we cannot deny that enterprises have the freedom to dispose of their own property.

Of course, we can actually find some counterexamples for the above reasons, so most of these statements are just excuses. The more fundamental reason actually comes from economic efficiency —

If large model enterprises are required to pay for all training materials from the very beginning, the manpower and material resources spent only on finding copyright holders, plus the licensing fees paid for each material, will be enough to stifle AI applications in the cradle. This is obviously not conducive to innovation.

For the overall benefit of society, we have to accept this kind of "fair use", even though the definition of "fair use" has always been very vague.

Photo | The Foundation for American Innovation

Therefore, if Anthropic had been willing to spend tens of millions of dollars to purchase legal books as the source of training materials at the very beginning, there would not have been much to criticize, and today's copyright lawsuit would most likely not have happened either.

Their mistake is that they chose to download pirated content instead.

Photo | Most of Anthropic's pirated training materials were downloaded from LibGen

Once pirated books are used as training materials, "fair use" is no longer sufficient to justify the whole process. It is like stealing a car to carry out public welfare activities, which cannot whitewash the act of car theft. Since Company A has left a handle on the data source, it can be predicted that if this lawsuit continues, it will most likely lead to a similar result:

The model training part constitutes fair use, but the pirated material library constitutes infringement, and the compensation that should be paid still has to be paid.

In addition, music publishers have put forward another reason this time: different from books that are generally digested, understood and then rephrased, the commercial value of lyrics largely comes from the display of the original text. And after their tests, Claude can indeed output complete or nearly complete lyrics under certain prompts.

In other words, at this point Claude is indeed distributing new copies to the outside world, and it forms competition with existing paid lyric services such as Musixmatch and LyricFind.

Photo | Claude later fixed the lyric output guardrail

If this point is confirmed to be true, the balance of victory will obviously tilt towards the music publishers.

Record Companies Can Always Find "Toll Booths"

Of course, if we only regard this incident as the just rights protection of music publishers, it would be a bit unfair to the negative reputation they have accumulated over many years.

For example, the aforementioned "paid lyrics" is a somewhat abstract business model for Chinese readers who are reading this article.

Because in China, thanks to the extensive cooperation between search engines and music publishers, as well as the high cost and low return of pursuing lyric copyrights, lyrics have largely become something that can be publicly displayed, accessed and discussed in our daily culture.

In almost any domestic AI product, you can openly search for the full lyrics of every song.

But abroad, due to the copyright network woven layer by layer by music publishers, lyrics have long been a kind of product that can be separately authorized and sold, and have spawned various "sub-landlords" that operate lyric licensing businesses. Correspondingly, all AI companies will strictly guard against their products being induced to output full lyrics and violate copyright laws.

The highly segmented copyright licensing system means that music publishers have long been selling far more than just a single song.

Listening to a song, covering it, putting it in a movie, putting it in a short video, displaying its lyrics, and even using it to train models now, can all be split into different licensing scenarios and priced separately.

Photo | Recording Arts Canada

Whenever technology creates a new way of usage, music publishers can quickly realize that: A new licensing pricing method may be created here.

From the era when iTunes and Spotify rewrote the digital music distribution method, to the short video era led by TikTok, music publishers have always benefited from the popularity created by platforms, while constantly using lawsuits and withdrawing music libraries as bargaining chips to demand licensing contracts, fight for higher licensing fees and longer cooperation terms.

Now it is just the turn of generative AI.

Suno and Udio, which focus on AI-generated music, were previously sued by large record groups, and have successively reached settlement agreements, preparing to launch new models trained with official music libraries, and allowing musicians to decide whether their works and sounds can be used for AI training.

Obviously, Universal, Sony and Warner are now aiming at Claude, wanting to bring it into this territory as well.

Music publishers, who have always been regarded as representatives of the "old world", have actually mastered the method of entering the "new world" proficiently:

First use lawsuits to stop the new order that is not under their control, and then implant themselves into the new business model through settlement agreements.

Photo | troveo

Therefore, what music publishers are targeting this time is far more than just the lyrics that Anthropic has downloaded.

They hope to take this opportunity to establish a previously non-existent market for "AI training lyric licensing".

Who Has the Right to "Share the Cake"?