Millions of books were "burned after reading" by Claude
A thick book is fed into a hydraulic paper cutter. The pressing plate lowers, the blade slices down, and the spine is neatly and cleanly sheared off. The pages that were originally connected turn into a neat stack of loose sheets.
With no additional context, you might even find this kind of video "oddly satisfying" and extremely stress-relieving.
But if you learn that AI companies are using this exact method to process physical books in batches, which may even include obscure and out-of-print titles, you will almost certainly no longer feel relaxed about it.
To source more training data, AI companies have expanded their search scope from the public internet and pirated e-books to physical books in the real world. Books published before 2022, which contain no AI-generated content and have not been processed by modern data poisoning tools, have become rare "high-quality assets".
Those obscure and out-of-print books that cannot be found online and have no electronic versions can provide content that web crawlers cannot capture, making them even more "top-tier assets".
The exposed "Project Panama" of Anthropic explicitly states that it does not want the outside world to know about this operation.
ISBNdb, a book database company, publicly claims that it can take orders to source books for AI companies, procuring 1,000 to 1 million physical books at a time from used bookstores, libraries and out-of-print book catalogs, and will hide the identities of buyers and the list of procured books through non-disclosure agreements.
AI companies are also fully aware that this practice is "unethical and unseemly".
"Project Panama"
The fact that AI companies have set their sights on physical books first came to light in a lawsuit involving Anthropic.
In August 2024, three writers sued Anthropic in court, accusing the company of using their works to train Claude without authorization.
To be honest, this kind of dispute has never stopped since OpenAI launched ChatGPT, and it is far from a new issue. There have already been many similar lawsuits.
The core of the controversy lies in whether it is an infringement for AI companies to crawl the entire human internet to train models? When an author's book is included in a training set, should the AI company obtain the author's consent and pay copyright royalties?
Therefore, in this lawsuit, physical books were not the focus at all in the beginning.
To make Claude more knowledgeable and capable of generating better content, Anthropic first scanned the entire public internet. However, the quality of content on web pages varies greatly, so books that have gone through author writing, editorial screening and publisher proofreading became their new target.
The most convenient method, of course, is to directly obtain e-books.
Court documents show that Anthropic once downloaded more than 7 million books from resource libraries such as Books3, LibGen and PiLiMi. A large part of these books came from pirated websites. The company did not delete these files after using them, but stored them in a long-term "central library" for continued use in future model training and research.
Subsequently, an internal project codenamed "Project Panama" was launched. In short, this plan refers to Anthropic's large-scale purchase, scanning of physical books to feed AI models.
But at that time, what the court really cared about was whether the way Anthropic obtained and used these contents was legal.
In June 2025, Judge William Alsup issued a ruling that was quite favorable to Anthropic.
In his view, when a large language model reads a book, its purpose is not to retell or sell the book to users, but to learn language, knowledge and expression patterns from it, and then generate new content. This use is highly "transformative", which is somewhat similar to the behavior of humans acquiring knowledge after reading a book and then creating new works, so it can be regarded as fair use.
As for the purchased physical books, Anthropic has paid for each of them. After scanning a book, the company destroys the corresponding physical copy and only retains a digital substitute internally, without making an extra copy for sale on the market. The judge therefore ruled that this part of the behavior did not infringe copyright.
However, the more than 7 million e-books downloaded from pirated resource libraries cannot be justified.
The judge held that although model training may be fair use, it cannot conversely prove that obtaining content through piracy is legal. How you use a book is one thing, but how you get the book in the first place is another.
In the U.S. legal system, precedents play a huge role. This judge's ruling has opened a path for AI companies: AI companies can learn from works created by humans, on the premise that they obtain these works through legal means first.
Therefore, it can be seen that in last year's Anthropic lawsuit, although its practice of purchasing and scanning physical books has been made public, the focus of the court and the public still remained on the long-discussed "legal boundary of using human knowledge to train AI".
It was not until two consecutive reports were released this year that people suddenly realized the severity of the problem.
Will Out-of-Print Physical Books Disappear Forever?
In January this year, The Washington Post dug up more details of "Project Panama" from more than 4,000 pages of newly unsealed court documents.
Anthropic spent tens of millions of dollars to buy millions of physical books in about a year. Suppliers cut off the spines, fed the loose pages into high-speed scanners, and finally sent the remaining paper to recycling companies. One single project plan alone planned to process 500,000 to 2 million books within half a year.
The scale is already staggering, and two sentences in Anthropic's internal documents are even more shocking:
"Project Panama is our operation to perform destructive scanning of all books in the world."
"We do not want the outside world to know that we are doing this."
When these two statements are put together, the whole thing becomes extremely problematic.
Anthropic has always emphasized AI safety, ethics and responsibility, but it secretly launched a secret operation to destroy millions of books.
Cutting book spines could previously be explained as an efficiency-oriented scanning method, but the statement of "not wanting the outside world to know" makes it hard not to suspect that Anthropic itself knows clearly that once this matter is made public, it will never be justifiable.
How many books has Anthropic secretly destroyed?
The Washington Post did not obtain the complete list of books, and the outside world has no idea how many of these books are widely available bestsellers and how many have been out of print. People are shocked by the industrialized scale of book destruction, but it is still difficult to assess what specific losses it has caused.
Six months later, a report from 404 Media focused public concern on out-of-print books.
The protagonist of the report is ISBNdb, a company that claims to own the world's largest book database. It publicly sells physical book procurement services to AI clients on its official website, claiming that it can procure 1,000 to 1 million books at a time from sources covering used bookstores, libraries and out-of-print book catalogs.
The company even takes confidentiality as a selling point: every project will sign a strict non-disclosure agreement, and the client's identity, procurement strategy and target book list will never be disclosed.
What is even more ironic is that ISBNdb itself knows this business cannot be brought to light. It directly reminds customers in its promotion: "AI companies destroying 2 million books" is not a news headline that can win public sympathy.
There is another interesting point: AI companies have a particular preference for physical books published before 2022 recently.
The reason is simple: after the generative AI boom, a large amount of AI-generated content has been mixed into the internet, and some people deliberately use data poisoning tools to interfere with model training. In contrast, physical books published before 2022 can almost certainly be confirmed to be written by humans, have gone through editing, proofreading and official publishing processes, and have not been processed by modern data poisoning tools.
AI companies themselves have created more and more online garbage, and finally turn around to look for old books that have not been polluted by AI.
This demand has already spread to the second-hand book market. A bookseller specializing in rare and low-circulation books told 404 Media that since April this year, the number of orders he processes per week has surged from around 20 to hundreds. The books purchased cover scattered themes and different languages, with almost no connection to each other, and their only common feature is that they all have an ISBN number.
This bookseller cannot confirm that the buyers are AI companies, but many of the books in his inventory are foreign-language books, obscure titles and out-of-print books.
On the one hand, he cleared out the inventory that had not been sold for years with these orders, but on the other hand, he felt very uncomfortable: "I don't like the final use of these books, and I don't want those rare books to be turned into paper pulp." As mentioned earlier, the process that AI companies use to scan physical books is very crude, which makes the second-hand booksellers whose business has improved also feel distressed.
From the internet to pirated e-books, and then to physical books that can be purchased in batches, AI companies have been extending their tentacles for training resources outward. Now, obscure books and even out-of-print books cannot escape this fate.
Out-of-print does not mean there is only one existing copy. There is currently no evidence to prove that the last remaining copy of any book in the world has been destroyed by an AI company.
The outside world cannot see the procurement lists, do not know which companies are buying books, nor know whether anyone has checked how many remaining copies of a book exist before it is sent to the cutter.
But the risks are obvious.
Extremely Limited Concessions
The 2025 ruling (at least in the United States) has given Anthropic a certain degree of legal protection.
The judge held that Anthropic legally purchases a physical book, scans it into a digital file, then destroys the physical original, and there is still only one copy in the end. The digital file is not sold to the public, nor does it increase the number of copies circulating in the market, which is equivalent to replacing one format with another, so it constitutes fair use.
According to this logic, cutting book spines has even become a part of proving the legality of the behavior. If the original physical book remains on the market while Anthropic holds an extra digital version, it may involve unauthorized reproduction; destroying the physical book and only retaining the digital file instead leads to lower legal risks.
James Grimmelmann, a professor of digital and information law at Cornell Tech, believes that Anthropic's move to abandon pirated book libraries and instead purchase and scan physical books is a "smart choice", which represents a more restrained and legally compliant practice.
Destroying a physical book is not a particularly sharp problem for a bestseller with hundreds of thousands of copies printed, as the loss of one copy will not affect other people's ability to read and purchase the book.
But for obscure books, out-of-print books and editions with special historical traces, this calculation obviously does not apply.
Dawn Albinger, President of the Australian and New Zealand Antiquarian Booksellers Association, pointed out that the value of antiquarian books lies not only in the printed text. The pages may have annotations left by people who experienced the related events. Signatures, inscriptions and information about previous collectors can prove the book's circulation history, and the rebound cover may even hide earlier manuscripts inside.
All this information is attached to that specific physical book. Even if the text on the pages has been scanned, once the original paper, binding and the relationship between all parts are destroyed, future generations will not be able to re-examine them.
In other words, AI companies obtain a digital file suitable for model training, but may destroy a physical object that cannot be completely restored through digital files.
After the controversy expanded, the technology industry soon made some concessions.
Elon Musk stated on X that he has asked the SpaceX AI team to store rare books in libraries and use more cumbersome non-destructive scanning methods instead of simply cutting off book spines. He did not disclose how many books SpaceX AI has purchased, nor did he explain what kind of books will be classified as "rare". This response is more like a show of stance.
ISBNdb made concessions even faster.
Nine days after the 404 Media report was released, ISBNdb deleted the physical book procurement page for AI companies, and removed all publicity related to confidential procurement and destructive scanning. The company later changed its statement, claiming that the relevant page was only a market demand test, the service had never been actually launched, and ISBNdb had never purchased, scanned or sold any book for AI training purposes.
Not long ago, it was still promoting the procurement of 1 million books at a time, but after public backlash, the entire business suddenly became "a test".
Whether this explanation is convincing or not, it at least shows that anonymously helping AI companies purchase physical books in large batches has become a reputational risk that enterprises are unwilling to take.
However, Musk's one promise and ISBNdb's deletion of several web pages cannot replace real formal regulations.
Who is responsible for judging whether a book is rare before scanning? Is the judgment based on the year of publication, edition and existing stock, or the signatures, annotations and binding of the book? Should AI companies disclose the list of books they purchase in large quantities? If a book is already out of print, should it only be scanned non-destructively and handed over to a library for preservation after scanning?
At present, there are no answers to these questions.
Ironically, ISBNdb previously repeatedly emphasized that the reason why physical books published before 2022 are precious is that they preserve human knowledge that has not been polluted by AI, and has been written by authors and screened by editors. The more AI companies recognize that this content is irreplaceable, the harder it is for them to justify why the physical carriers that hold this content can be treated as disposable consumables.
This article is from the WeChat official account "Alpha AI", written by Xiao Jinya, and authorized for release by 36Kr.