HomeArticle

The U.S. Department of Justice has thrown its weight behind OpenAI: Can AI training be conducted without paying copyright royalties?

知产力2026-09-03 14:53
The copyright disputes over AI training in the United States offer important implications for China's development.

On September 1 local time, the U.S. Department of Justice filed a Statement of Interest of the United States Government in the consolidated lawsuit filed by publishers including *The New York Times* and authors against OpenAI and Microsoft, with the U.S. District Court for the Southern District of New York.

The core position of the document is very clear: using copyrighted works to train large language models is usually highly transformative in itself; the court should not issue a broad ruling that generally identifies unlicensed AI training as copyright infringement.

The U.S. Department of Justice even proposed that if the fair use doctrine is incorrectly narrowed, it will not only hinder technological innovation and scientific progress, but also undermine the U.S. economic competitiveness and national security, putting U.S. AI enterprises at a disadvantage in global competition.

*The New York Times* immediately responded that the U.S. government is sacrificing the interests of creators to stand for a few trillion-dollar AI companies; AI enterprises should pay reasonable remuneration for the content that supports their products.

On one side are the AI industry, national security and global competition, and on the other side are news production, copyright revenue and the creative ecosystem. Whether AI training should be paid for is no longer just a copyright law exam question, but is becoming a policy issue concerning what AI development path the United States will choose.

I. The statement of the Department of Justice does not mean that OpenAI has won the lawsuit

This document is not a court judgment. The U.S. Department of Justice submitted the "Statement of Interest" in accordance with Title 28, Section 517 of the U.S. Code, for the purpose of stating to the court the position of the U.S. government on the public interest and legal issues involved in this case. It can influence the judge's judgment, but it is not automatically binding on the court, nor does it mean that the case of *The New York Times* v. OpenAI has reached a conclusion.

The Department of Justice also specifically stated that the government has not authorized, consented to or endorsed any specific conduct alleged in this case.

Therefore, a more accurate expression is: the U.S. Department of Justice supports OpenAI's core defense that fair use applies to AI training, but whether OpenAI constitutes infringement in specific links of data acquisition, model training and content output still needs to be judged by the court based on evidence.

Nevertheless, the significance of this opinion remains significant.

According to public reports, this is the first time that the U.S. government has formally expressed its position to the court on the copyright nature of large model training amid the wave of concentrated lawsuits filed by copyright holders against AI enterprises. The interest game that previously mainly occurred among enterprises, publishers and creators has thus obtained a clear national policy orientation for the first time.

The U.S. government is not resolving a lawsuit for OpenAI, but trying to gain more training space for the entire U.S. AI industry.

II. Why does the Department of Justice believe that AI training falls under fair use?

The first key argument of the U.S. Department of Justice is to divide large model development into three links:

First, acquiring and storing data; second, copying works for model training; third, generating outputs based on user instructions.

The Department of Justice believes that these three links may give rise to different copyright issues respectively. We cannot conclude that all previous training acts are automatically illegal just because a model occasionally outputs content similar to the original work.

At the training stage, the model may indeed copy complete works. However, these works are not directly provided to the public as articles, novels or images, but are converted into digital information for analyzing the statistical relationship between words, grammar, knowledge and expressions, so that the model can generate new responses to various questions.

The purpose of the original work is to provide specific content to readers; the purpose of the training act is to build a general tool that can understand and generate language.

It is this difference in use that makes the U.S. Department of Justice refer to large model training as "highly transformative use".

This line of thinking continues part of the judgments made by U.S. courts in the 2025 hearings on the training data cases of Anthropic and Meta. The preliminary tendency emerging in U.S. courts is not that "AI enterprises can use copyrighted works at will", but to judge the issue separately:

Training can be transformative, but the data source may not be legal; the model can learn from works, but the output cannot automatically copy the works.

What the U.S. Department of Justice intends to strengthen this time is exactly this legal path of "judging by separate links".

III. Why may complete copying of works still not constitute infringement?

The most direct query from *The New York Times* and other right holders is: if AI training needs to copy full articles and entire books, why is there no need to obtain permission from the authors?

The U.S. Department of Justice's response is that fair use does not only depend on "how much is copied", but also on why the copying is done, and what the copied result provides to the public.

For example, to build a full-text search system, it may be technically necessary to scan an entire book completely; to train a language model, it may also be necessary to input the full article into the model. But if what is finally provided to the public is not the complete work, and does not form a substantial substitute for the original work, complete copying itself does not necessarily exclude fair use.

The Department of Justice believes that copying at the training stage will not directly present the original work to the public. Even if the model may produce memorized output in very rare cases, such output should be judged separately for specific instances, and the entire training process should not be negated by a small number of abnormal outputs.

This set of logic tries to establish a very important boundary:

Copyright controls the copying and dissemination of the expression of works, not the facts, knowledge, language rules and general creation methods contained in the works.

If the model only learns word relationships and factual information from a large number of articles and generates new content that is not substantially similar to the original text, the Department of Justice believes that this kind of market competition is not a substitute for works in the sense of copyright law.

But if the model can reproduce important paragraphs of the article, or users can obtain substantially the same content without accessing the original media, the situation will change. AI can learn knowledge, which does not mean it can deliver the original work.

IV. The real dispute is whether the "training licensing market" should exist

The most difficult part of this case to resolve is not whether the model has copied articles, but how far the market of copyright holders should extend.

Right holders such as *The New York Times* claim that news organizations have begun to sign content licensing agreements with AI enterprises. If OpenAI uses their articles to train models without permission, it will deprive news organizations of the licensing revenue they deserve, and use such content to develop products that compete with news media.

The U.S. Department of Justice argues that the mere fact that a copyright holder is willing to sell a license does not automatically mean that the law must recognize this licensing market.

Otherwise, as long as a copyright holder requests payment for an act that might originally fall under fair use, it can prove market harm by "not receiving licensing fees", and then use market harm to negate fair use in reverse. This will form circular reasoning.

The Department of Justice further proposes that if all training materials must be paid for one by one, only the most capital-rich tech enterprises can afford the cost, which will instead consolidate the advantages of large AI companies and prevent small and medium-sized model enterprises from entering the market.

At the same time, large traditional media with massive historical content will obtain the most licensing revenue, while independent media and new creators may not be able to fairly participate in the distribution.

This is the most noteworthy change in this opinion.

In the past, AI enterprises mainly argued from the technical level that "training is not copying works, but learning rules"; now the U.S. government has added an industrial policy argument: mandatory training licensing may simultaneously strengthen large tech companies and traditional content giants, raising the entry threshold of the AI market.

As a result, copyright licensing is no longer just a matter of whether creators can get remuneration, but is also being measured in the context of the competitive structure of the AI market.

V. Training can be free of charge, which does not mean the content has no value

The U.S. government supports the application of fair use to AI training. Does this mean that AI enterprises no longer need to pay for content in the future?

The answer is no.

The Department of Justice's opinion itself acknowledges that regardless of whether training constitutes fair use, AI enterprises can still sign commercial licensing agreements with media and content institutions to obtain real-time content, paywall content, professional databases and other high-quality information.

This reveals the possible structural changes that will occur in the future content licensing market.

What AI enterprises purchase may not only be "legal licenses for copying works", but also include:

First, legal, stable and sustainable data access; second, real-time updated professional content with clear structure; third, credible data that can mark sources and trace evidence; fourth, high-quality corpora for retrieval augmentation, model evaluation and professional responses; fifth, certainty to reduce litigation, compliance and brand risks.

In other words, even if U.S. courts finally determine that general model training falls under fair use, it does not mean that all content has lost its commercial value.

Public web pages, pirated materials, paid databases and real-time news will not have the same legal status and market value just because they can all be converted into tokens.

The focus of copyright charging may gradually shift from "whether the model has seen the work" to "where the data is obtained, whether it can be called continuously, whether the output can be traced, whether it replaces the original work, and whether it can support reliable commercial applications".

Copyright will not disappear, but the charging nodes may change.

VI. China cannot directly copy the U.S. solution

The opinion of the U.S. Department of Justice has an important impact on the global AI industry, but its conclusion cannot be directly transplanted to China. The U.S. fair use system has strong openness, and courts can make comprehensive judgments around the purpose of use, the nature of the work, the amount of use and the market impact.

Article 24 of China's Copyright Law mainly lists specific circumstances where works can be used without permission and without remuneration, and requires that the normal use of the work shall not be affected, nor shall the legitimate rights and interests of the copyright holder be unreasonably damaged. At present, China has not established a clear text and data mining exception for commercial large model training.

Therefore, under Chinese law, whether model training involves the right of reproduction, whether the data source is legal, whether the platform has obtained authorization, and whether the relevant use can apply to the limitation of rights still need to be judged in combination with specific acts, and the U.S. "fair use" cannot be simply cited as a basis for exemption.

However, this U.S. dispute still provides an important reminder for China:

If every training act is required to obtain authorization one by one, the transaction cost may be too high to implement; if commercial training is completely liberalized, the continuous supply of professional content and original works may be weakened.

What really needs to be established is not a simple rule of "all free" or "all paid", but a layered system formed according to data sources, work types, model uses and output risks.

Public information, factual data and scientific research analysis should be treated differently from pirated works, paid content and professional databases; the basic training of general-purpose models should not apply exactly the same rules as industry models, real-time calls and alternative outputs.

Zhichanli's Judgment

The U.S. government standing on OpenAI's side does not mean that the AI copyright dispute has ended.

The signal it really releases is: the United States is trying to establish a broad fair use space for general model training, and leave the main responsibilities to the links of illegal data acquisition, work reproduction, market substitution and specific output.

Therefore, for the question raised in the title, the answer may be:

AI enterprises may not have to pay copyright fees one by one for every pure training act, but this by no means means that they can obtain all content for free, let alone copy and substitute creators' works without remuneration.

The most valuable content in the future will not only be the content that can be "digested" by the model, but the content with legal sources, continuous updates, clear structure, verifiability and safe callability.

What the content industry needs to prove next is that although knowledge can be learned, high-quality, verifiable and continuously updated knowledge has never been free.

This article is from the WeChat official account "Zhichanli" (ID: zhichanli), author: Shawn/MCP, published with authorization from 36Kr.