HomeArticle

"Annotated books are strictly prohibited from being used for AI training!" Huaxia Publishing House: It is difficult to safeguard legitimate rights, but we must take a firm stance.

36氪的朋友们2026-08-05 15:41
The biggest difficulty in rights protection lies in evidence collection, and it is usually very difficult for the party safeguarding legitimate rights and interests to obtain the internal training datasets of AI enterprises.

Recently, readers have noticed that the copyright page of books including *New Translation of Huangting Jing and Yinfu Jing* published by Huaxia Publishing House is clearly printed with a warning line: "The content of this book is prohibited from being used for artificial intelligence training, and violators will be held accountable." This book is annotated and translated by Liu Lianpeng and Gu Baotian, published and distributed by Huaxia Publishing House Co., Ltd., and the property rights of the works belong to San Min Book Co., Ltd.

On August 4, according to Red Star News, the relevant person in charge of Huaxia Publishing House stated that when introducing some foreign books, the copyright holder will add clauses in the copyright contract to prohibit the text of the books from being used for artificial intelligence training. "As the performing party, we need to give an account to the rights holder, so we will add this sentence."

However, the person in charge said that if the content of the book is used for artificial intelligence training without authorization, it is actually very difficult to safeguard rights. The publisher's marking is mainly a statement to make people realize that books are copyrighted.

At present, many paper book publishers cannot rely on technical means to automatically block scanning and entry behaviors for the time being. Unlike websites that can set crawler protocols, physical books do not have online protection mechanisms. For organized data collection teams, the warning text is more like a "gentleman's agreement". For ordinary publishing institutions and creators, if their book content is illegally collected for AI training, the feasible rights protection paths are still limited.

Some legal professionals suggest that after discovering clues of infringement, the ownership evidence such as the copyright page of the book and the sample of the published work should be fully preserved first. Once a company is found to scan books in batches, a written letter can be sent to demand that the collection behavior be stopped. However, the biggest difficulty lies in producing evidence. Ordinary people can hardly obtain the internal training data sets of AI enterprises. At this stage, the industry relies more on prior agreements and continuous public disclosure of copyright claims to reduce the risk of infringement.

In fact, the copyright game between the overseas publishing industry and AI enterprises has long been staged. In early 2024, artificial intelligence startup Anthropic launched an internal secret project codenamed Project Panama, whose core goal was to obtain "all the books in the world". In the early stage, Anthropic relied on pirated e-books such as Books3 and shadow libraries to obtain training texts, triggering a collective copyright lawsuit from writers.

Court documents submitted by Anthropic last September showed that the company would pay at least 1.5 billion US dollars to settle the above lawsuit. According to the settlement agreement, Anthropic will pay approximately 3000 US dollars in compensation to the authors or publishers of about 500,000 books included in the settlement scope.

To avoid the legal risk of piracy, the company turned to purchase second-hand physical books in batches to collect high-quality training corpora, and directly destroyed the original books after the entry was completed. About 2 million paper books were damaged in the data collection process. After the incident was exposed, the controversy of "destroying books for AI training" swept the global publishing industry, and a large number of publishers began to re-examine the data security of physical texts.

Compared with the public Internet texts full of fragmented remarks, repeated information and wrong content, formally published books have been checked layer by layer, with coherent writing and complete logic, which can help large models master systematic expression and long-text reasoning capabilities. A large number of ancient books, professional works and out-of-print documents do not have complete electronic versions circulating on the Internet, and their full content can only be obtained by scanning physical books. Multiple types of publications can also enrich the knowledge boundary of the model and balance various capabilities such as written expression, narration and theoretical interpretation.

The extremely high training value makes books an indispensable corpus source for large model training. This also explains why some AI enterprises spare no cost to collect paper books in batches for digitalization. However, the ensuing copyright conflicts are forcing the publishing industry to actively delineate the boundaries of rights.

The academic circle is also taking action. This May, Elsevier, the leading publishing group that owns journals including *The Lancet* and *Cell*, jointly sued Meta with four other publishers, accusing it of obtaining a huge amount of copyrighted books and journal content without permission for the training of the Llama series large models.

The indictment shows that in order to seize the opportunity in the AI arms race, Meta not only used web scraping datasets containing billions of web pages, but also downloaded and disseminated millions of copyrighted books and paid academic journal articles from pirated websites. Meta is also accused of deleting copyright notices and author information in the works to cover up the data source.

In response, Meta argued that using copyrighted content to train AI is transformative use, which can apply the fair use rule in the US Copyright Law, and it will actively respond to the lawsuit.

Last March, major French publishers and writers associations also filed a lawsuit against Meta, accusing Meta of using copyrighted content on a large scale without authorization to train its AI models.

Up to now, there is still no unified and clear judicial conclusion around the world. Courts in different regions have issued completely different judgments for similar cases.

In March this year, according to 21st Century Business Herald, Tao Kaiyuan, Vice President of the Supreme People's Court of China, said when talking about the challenges brought by generative artificial intelligence to intellectual property judicial adjudication that judges often cannot make judgments by simply "matching" the current laws, but they must answer the "judicial questions" brought by the rapid development of artificial intelligence technology.

In Tao Kaiyuan's view, it is necessary to seek a balance between the development of the artificial intelligence industry and the protection of the legitimate rights and interests of copyright owners. She revealed that China is currently actively promoting the drafting and formulation of judicial policy documents involving artificial intelligence and data property rights, striving to provide a clear judicial boundary for the application of new technologies.

This article is from the WeChat Official Account "Jiemian News", written by Song Jiannan, and authorized for release by 36Kr.