HomeArticle

How did Anthropic destroy books to train large models?

明亮公司2026-09-10 18:42
Anthropic believes that books are the "highest-quality source of training data" for its large models.

On September 8, *The New Yorker* published an article titled "Destroying Books to Build a Mind", discussing the controversy over AI company Anthropic's large-scale purchase and "destructive scanning" of physical books to train its Claude model.

In the article, the author starts with a phenomenon in early 2026, when multiple second-hand booksellers in the United States and Canada suddenly received a large number of orders for niche, obscure academic books, and the buyers often appeared under the name of mysterious companies. It was later discovered that these books were sent to industrial scanning facilities, where their bindings were removed, pages were cut, scanned into digital texts, and the physical copies were discarded. Anthropic referred to the project as the Panama Project in legal documents, with the goal of accessing "all the books in the world".

The article points out that destructive scanning itself is not uncommon, and libraries also dismantle books for digitization. What really raises concerns is whether AI companies are destroying rare and precious books, and whether the original text, author background and reading context are erased after these books are converted into training data. Multiple booksellers and experts believe that most of the books purchased by Anthropic are not ancient books or rare editions, but obscure academic books, professional books and low-circulation books published since the 1970s. Such books might otherwise become obsolete, and could potentially have a greater impact after being incorporated into the model.

At the legal level, the judge ruled that using legally obtained books to train AI constitutes fair use, and format conversion to digital copies is also valid.

However, the article emphasizes that the deeper issue is not the "book destruction" itself, but transparency and cultural significance. Anthropic's so-called "research library" is not designed for people to read, but to extract data and expand the model's memory. Researchers and users cannot know which books the model has absorbed, nor can they trace the source of information.

The author concludes in the article that a book is more than a collection of sentences. Real reading involves time, sequence, context, imagination and the relationship between people and text. Large language models convert literature and knowledge into recombinable data, which may cause a more subtle erasure than physical book destruction. Nevertheless, the author also believes that there will always be people who prefer summaries or rewritten content generated by the model, while others will continue to pursue original texts and real reading experiences.

It is worth noting that Anthropic identifies books as the "highest quality source of training data" for its large language model. For books that have not been digitized, there is no public information about the practice of Chinese large model companies regarding whether they can enter the training dataset and how they can do so up to now.

Anthropic confidentially submitted its draft IPO filing on June 1. It was previously expected to be publicly disclosed in early September, but the disclosure time has now been delayed at least until late September. The latest publicly disclosed equity financing of Anthropic was in late May, with a post-investment valuation of 965 billion US dollars, and its ARR exceeded 65 billion US dollars by the end of July. The market expects the company's market capitalization to exceed 2 trillion US dollars after listing.

Many American investors have publicly complained about the negative attitude of mainstream American media towards AI companies, and believe that the American public's pessimistic expectations for AI (such as refusing to build data centers locally) are closely related to the attitude of mainstream media.

The following is the content of the article compiled by Suchbright:

Destroying Books to Build a Mind

Anthropic is trying to "destructively scan all the books in the world". How worried should we be?

Author: Francesca Mancino

In early 2026, the owner of James Payne, Books and Prints, an old bookstore in Brooklyn, began to receive a batch of order requests. He later described these orders as "outrageously bold and outrageously weird".

The books themselves are strange: for example, a 1998 so-called legal survival guide *How to Win in New York Small Claims Court*, and a 2006 academic work *The Culture of Glass Architecture*. The ordering method is also strange. Booksellers like Payne usually list their inventories on multiple websites, such as AbeBooks, Alibris, Amazon and Biblio. Among these platforms, Alibris is usually regarded as the most deserted revenue channel. But suddenly, Payne's sales on Alibris changed from less than 10 books a year to processing 10-book orders at a time, with an average price of 75 to 80 US dollars per book. Other booksellers across the United States have reported similar situations.

Sylvia Petras is the owner of Leaf and Stone Books in Toronto. At first, when she saw American booksellers posting news of surging sales online, she was quite envious. Then, orders began to flow into her store one after another, and that envy quickly turned into confusion. Petras has a large collection of printed books from the 15th to 17th centuries, as well as rare books and academic books; some books with weird themes, such as books about sewage treatment plants, suddenly flew off the shelves.

Like other booksellers overwhelmed by orders for niche books, Petras noticed that these orders were not placed in personal names, but from mysterious limited liability companies: Green Parrot Project and Red Sparrow Project. A bookseller who requested anonymity decided to install a GPS tracker on a purchased book before sending it to the mysterious buyer. After researching several options, she chose a SmartCard device that looks like a credit card. She put it in a small envelope, then fixed the envelope to the inside of the back cover of a hardcover book, behind the dust jacket. She then tracked the book's whereabouts online: it departed from Grandview Heights, Ohio, and arrived at an industrial park in Addison, Illinois. There is a company called ARC Document Solutions that operates industrial-grade scanning facilities. In a promotional video, a spokesperson for ARC explained that many of the company's sites operate around the clock, with more than 100 across the country. He also mentioned that a client hired ARC to scan "boxes the size of a football field stacked four stories high every month".

Anthropic is the artificial intelligence company behind the Claude series of large language models. It is trying to acquire as many physical books as possible. We know this from Anthropic's own legal documents, which were unsealed in a copyright lawsuit filed against the company in 2024. After Anthropic identified books as the "highest quality source of training data" for its large language model, it launched a secret project called the Panama Project. The concise definition of this project in court documents is: "Our effort to destructively scan all the books in the world".

The meaning of "destructive scanning" is very precise: according to court documents, Anthropic or its affiliates "remove the binding of physical books, cut the pages into workable sizes, then scan these pages, and discard each physical copy while creating a digital copy".

This process uses a hydraulic cutter, the so-called "book guillotine". It sounds creepy, but this practice is actually quite common. The document preservation departments of universities often remove the bindings of books to make scanning easier. Of course, destroying a book is not the only way to scan it. When digitizing texts, Princeton University Library specially preserves physical copies, even if it makes the process less efficient. Meredith Martin, a professor of English at Princeton University and director of the Center for Digital Humanities, told me: "It is entirely possible to capture images without flattening the book and preserve the integrity of the book, but it costs more."

However, Martin and other experts also emphasized that old books have always been discarded and then destroyed, not necessarily for academic reasons. Some libraries just want to clean up their collections. Donation is certainly more acceptable, but it is not easy to dispose of an outdated instruction manual or textbook, or several tattered copies of the same mass-market paperback. Some books will not even be accepted by prison libraries, which usually refuse hardcover books or books in poor condition. As a result, these books may end up in landfills; or they may be disassembled, pulped for recycling, or shredded depending on their binding.

Claire Sewell, an academic librarian in Houston, once wrote an article on Medium defending libraries' "weeding" of bookshelves. She admitted that "seeing an entire dumpster full of books seems to go against everything that a library as a repository of knowledge should represent." However, "old, outdated, damaged, or simply low-circulation books must be regularly weeded out so that we can make room for the new books you actually want to borrow."

What makes Anthropic's project controversial is not the fear that it will destroy a water-damaged copy of *The Secret*, but the fear that it will destroy truly rare and valuable books. To meet its voracious demand for billions of pages of content, Anthropic initially placed large orders with wholesalers, and then turned to individual booksellers and second-hand booksellers to fill the gap. The company's interest in obscure texts has led to headlines saying that AI companies are buying and destroying "rare old books" or "antique books". These claims usually conjure up images of first editions.

With the rise of AI companies like Anthropic, many legitimate ethical concerns have indeed emerged. But the destruction of precious books is not necessarily one of them. A spokesperson for Anthropic wrote in a statement: "None of our data acquisition projects purchase and destroy rare or ancient books." The term "ancient books" here refers to books that are at least 100 years old, whether rare or not. The Antiquarian Booksellers Association of America said no member has reported that such books were sold to AI companies. None of the individual booksellers I interviewed sold ancient books to buyers who looked like AI companies. Instead, most of the books they sold had ISBNs, that is, unique numerical identification codes; ISBNs were not used until 1970.

Petras, the Toronto bookstore owner, said she is generally willing to sell books to buyers with hidden identities, but she will never sell books with interesting marginal notes or important provenance background to them. Joyce Kosofsky, one of the owners of Boston's Brattle Book Shop, said all the books she sold had multiple copies. "We are probably the cheapest one," she guessed. She believes that in general, books are no different from any other salable commodity: "It's like you walk into a clothing store and buy a pair of jeans, that's your jeans. You can wear them, decorate them, give them away. No one will follow you and ask, 'What are you going to do with your jeans?'"

Since the books Anthropic bought are "very likely not very rare", Princeton professor Martin said the company's use of the book guillotine should not be regarded as an inflammatory issue. But if these books are not rare, what exactly are they?

Booksellers shared more than 600 titles that they believed had been sold to AI companies. I sent this list to Melanie Walsh, an assistant professor at the University of Washington Information School, to process digitally. These texts are characterized by obscure themes and low popularity. Some books have very small print runs, usually 1,000 copies or less, to meet the needs of the real market. Among the top ten publishers, eight are university presses; about one-third of the entire sample comes from academic presses. Most of the books were published between the 1970s and 2010s, covering history, biography, fiction, poetry, literary criticism, law and social sciences.

Walsh concluded: "From this sample, it seems that AI companies may be interested in training models with a wide range of peer-reviewed academic research." This is consistent with the findings of Anthropic research scientist Mycal Tucker when organizing AI model training data. Tucker said in written testimony: "Non-fiction works tend to be more valuable than fiction works." He added that non-fiction book data helps the model perform well in very different disciplines such as philosophy and astronomy.

Walsh said that for some highly specialized non-fiction works, such as a 1972 pamphlet *Utilization of Urban Sewage Sludge* that Petras recently sold, "one can make the argument that these obscure academic books may have a greater impact as part of the Claude model than they would have existed in other ways." After all, they were already becoming obsolete.

Anthropic's effort to access "all the books in the world" has also triggered legal challenges. In 2025, the company agreed to pay 1.5 billion US dollars to settle a class action lawsuit filed by a group of authors who accused Anthropic of copyright infringement, specifically using their books for AI training without permission.

Judge William Alsup, who presided over the case, condemned some of Anthropic's behaviors, such as downloading more than 7 million pirated books and keeping these files as "permanent universal resources", even if they were not used to train Claude. Alsup wrote in the ruling: "Anthropic seems to think that because some of the works it copies are sometimes used to train large language models, Anthropic has the right to take all the works in the world for free and keep them forever without further accountability."

But according to Alsup's judgment, the broader practice of using books to train AI models constitutes fair use. He wrote: "An author cannot justifiably exclude anyone from using their works for training or learning." He also said: "Everyone will also read the text and then write new text. They may need to pay when they originally obtain the text. But it would be unthinkable to require anyone to pay specifically for the use of a book every time they read it, every time they recall it from memory, and every time they later write in a new way and draw on it."

Brandon Butler, a copyright lawyer and executive director of Re Coalition, described AI training as "the fairest use in the history of copyright law". Re Coalition is an advocacy organization that supports both creators of proprietary works and their users. Butler believes that this is because AI training "takes from the existing culture only what actually belongs to all of us: facts, ideas, language and grammatical elements, and makes these things easier for everyone to use."

If a large language model verbatim repeats memorized material or conveys protected expressions, copyright may be infringed. But they usually convey unprotected information, and the way of expression is different from the wording of the original source. Butler said that in this sense, there is no market competition between AI companies and authors, because if a person sees a paraphrased excerpt of a book, they still have the motivation to buy that book. Of course, this issue is complicated by the fact that some large language models occasionally "memorize" almost entire novels, including Meta's Llama.

According to Judge Alsup's judgment, Anthropic's destruction of its legally obtained physical books is also within its rights. He wrote: "Anthropic purchased millions of physical books to 'build a research library'. It destroys each physical book and replaces it with a digital copy for use in its library; these copies are not shared or sold outside the company." Since this format conversion is for convenience, that is, easier storage and search, and does not increase the number of copies of the original book, it constitutes fair use.

But leaving legality aside, if Anthropic's understanding of the term "research library" is so different from ours that it makes the description almost a misnomer, we will encounter difficulties even at the linguistic level. Butler explained: "They are not building a library, because maybe one day someone will want to read these books. They have no interest in helping other people read these books. What they really want is data."

Their digital collections are built for the sole purpose of expanding the "memory" of Anthropic's large language model and honing its writing ability.

As a result, even if these models do successfully preserve literary materials that might otherwise be lost over time or due to declining interest, the source will always be in the shadows. Researchers cannot access this book data, and users cannot control or observe how information is provided to them. Mike Furlough, executive director of HathiTrust, told me that we have the right to demand that AI companies be more transparent about how their datasets are created, even though their mission is not to distribute books, and disclosing this data may sacrifice their competitive advantage.

Libraries and academic researchers are required to adhere to much stricter standards. Furlough said, "An academic researcher who wants to build a small language model will be required to disclose the complete bibliography of what they have used." He continued: "This is just common academic practice, because you want other people to be able to replicate it, or understand what was used in it." This is basically true for any reliable tool that ordinary people use for research; even Wikipedia includes citations. Anthropic's goal of gathering "all the books in the world" could have been almost exciting; the problem is that this hypothetical written knowledge base will only exist in the digital depths of Claude.

Historian Leah Price wrote in a 2019 non-fiction history of the book medium: "Every generation rewrites the epitaph of the book. The only thing that changes is who the murderer is." In the 19th century, libraries were suspected of being the murderers. In a pamphlet printed in London