HomeArticle

Are AI Giants the Culprits of "Book Burning"? Elon Musk Quickly Distanced Himself: He Has Required SpaceX AI to Conduct Lossless Scanning of Books

36氪的朋友们2026-07-30 09:16
Copyright disputes arise as AI enterprises purchase books to obtain training corpora, and the publishing industry is expected to benefit from the related development.

As AI-generated content floods the Internet like a tide, manually created "clean corpora" have suddenly become a scarce resource. Against this trend, AI giants are turning their sights to printed books.

Recently, a report from the independent media outlet 404 MediaX pointed out that AI companies are purchasing rare books published before 2022 in batches, using hydraulic cutters to remove the book spines and scan them, so as to obtain clean data that has not been "polluted" by AI (books published before 2022 do not contain AI-generated content).

To prevent the data from flowing back to the market, these books are often crushed and destroyed afterwards. The report noted that Anthropic has destroyed millions of physical books, even including extremely rare out-of-print ancient books with very few existing copies.

It is reported that this operation is called "Project Panama" by Anthropic, and has been ruled legal by the U.S. federal court. The court determined that the model of "legally purchasing printed books and then performing destructive scanning" falls under fair use.

As soon as the report was released, related public opinion quickly fermented on social media: The X platform account Hedgie received a reply from Elon Musk when relaying this incident. The latter said: "I have asked the SpaceX AI team to preserve all precious books in the library and adopt non-destructive scanning methods to avoid damaging the book spines."

Hedgie replied to this: "Thank you to Musk's team for doing this. Although the scanning speed is a bit slower, you can still get training data. If SpaceX AI can stick to this practice, other companies in the industry should also take it as the standard."

No matter what method is used to obtain data, various signs indicate that books have become one of the important sources of AI training corpora.

As mentioned in the above report, ISBNdb is a company that provides large-volume book procurement services. It claims to have "the world's largest book database" and has transformed into a "book data supplier" for AI enterprises in recent years. The Washington Post found that Anthropic has established cooperation with Better World Books, which is a platform for libraries, retailers and individuals to sell used books.

In this context, frequent legal disputes have broken out between overseas publishing media groups and AI enterprises:

On July 22, News Corp formally filed a counterclaim in the Federal Court of Oakland, California, accusing the American search engine company Brave Software of unauthorized large-scale scraping and reselling of articles from its media outlets including *The Wall Street Journal* and *New York Post* for artificial intelligence model training.

On July 14, a group of major publishers filed a lawsuit against Google, accusing the company of illegally using millions of copyrighted books to help build its Gemini artificial intelligence model, calling it "one of the worst copyright infringements in history". Google internal sources pointed out that if the text provided by publishers is used for Google Play Books, the company may face "potential fines of 10 billion to 100 billion US dollars".

The research report of CITIC Securities pointed out that the digital operation of cultural IP is expected to bring an incremental space of about 70 billion yuan to the publishing industry in 2026. Large model training requires a large amount of high-quality corpus data, and the content assets owned by publishing enterprises can provide rich high-quality language materials for the model. Zhongtai Securities stated that relevant publishing companies have clear ownership of content copyrights and obvious advantages in vertical content, and are expected to cross multiple technology cycles such as the Internet and AI, and become the underlying core assets for sustainable development.

This article is from the WeChat official account "CLSA AI Daily", author: Zhang Zhen, released with authorization from 36Kr.