"AI Book Burning" is actually a misunderstanding.
In Silicon Valley, the "AI book burning" saga has only grown more absurd and out of control.
In a latest report, media outlets partnered with secondhand booksellers to slip an AirTag into a batch of books ordered by a mysterious buyer, tracked the entire journey, and finally traced the shipment to Amazon's facility in Las Vegas.
Factory employees said this is a project codenamed VGT3 at Amazon, which is specifically dedicated to the work of slicing books for scanning.
A secondhand book has somehow followed a plot straight out of a crime investigation drama.
At this stage of development, the whole incident has become somewhat surreal.
In the past month, reports on "AI book burning" have emerged in endlessly, but there are actually no new facts: AI companies are purchasing large quantities of old physical books, cutting off their spines, scanning them at high speed, digitizing the content and then disposing of the original books. When Anthropic's "Project Panama" was exposed earlier this year, the most striking elements including millions of books, tens of millions of dollars in procurement and destructive scanning were already laid out in public.
After that, global secondhand booksellers spotted abnormal orders, ISBNdb was exposed to specifically source physical books for AI companies, and now the AirTag has tracked all the way into Amazon's warehouse. These developments have repeatedly confirmed one thing: yes, this is really happening, and more than one party is doing it.
It is obvious that the facts remain the same, but public sentiment around "AI book burning" has become increasingly intense.
"AI companies are devouring books", "Human civilization is being fed to machines", "They won't even let out-of-print books go" — countless people have left panicked comments on social media, and videos and photos of book slicing are particularly stirring public sentiment.
To be fair, there has been a significant misalignment between the known facts and the meaning the outside world has assigned to the incident as it unfolds today.
There is no need to cover for AI companies on the fact that they are indeed destroying books.
But how many books have they actually destroyed? Does destroying physical books equal destroying knowledge? Is the full responsibility for out-of-print books being bought and sliced really on AI companies?
It is time to break some myths.
01
Let's put the conclusion here first: A large part of people's current anger at "AI destroying books" points to the symbolic meaning of this incident, rather than the actual harm that has already occurred.
Of course, the two are easily mixed together. After all, headlines like "AI companies are buying and destroying antique books" and "AI companies shred millions of books after scanning" are really scary, and the book slicing process is extremely visually impactful. It is hard to find a more suitable scene to show that "AI is devouring human civilization".
But if we pull the camera back a little, the situation is not as apocalyptic as it seems in the footage.
The largest known project so far belongs to Anthropic. After launching "Project Panama" in 2024 with the goal of scanning all books in the world, it spent tens of millions of dollars in about a year, purchased and destructively scanned millions of physical books, sometimes buying tens of thousands of books at a time.
But the figure of "millions of books" is actually not as exaggerated as imagined when placed in the entire secondhand book market.
World of Books, just one large secondhand bookseller, sold 31 million books globally in 2024 alone, with a long-term inventory of about 7 million books and about 2 million types of books listed on its website all year round.
In other words, Anthropic's processing volume is only equivalent to a few months of business for a large secondhand bookseller at its normal circulation scale.
So it is true that "AI companies have destroyed millions of books after scanning", but it is far from the conclusion that "AI is eating up all secondhand books".
Moreover, there is another issue that is easily given extra emotional meaning that needs to be clarified: the disappearance of physical books does not equal the disappearance of the knowledge contained in the books.
Anthropic bought these books precisely to preserve their content. In the copyright lawsuit previously filed by writer Andrea Bartz and others against Anthropic, court documents show that after each physical book is disassembled and scanned, it will be turned into a PDF containing full-page images and machine-readable text, and added to Anthropic's own digital library.
If we only discuss the behavior of "buying a book, slicing it, scanning it, and then destroying the original", it is almost the opposite action of the historical "book burning".
Book burning is intended to make content disappear, while AI companies disassemble books specifically to preserve the content. Once used for model training, the content will be turned into new dynamic value. Judges have also expressed similar views: the content in books used for model training is transformed into new capabilities during the training process.
Those sliced physical books are certainly distressing to look at, but human knowledge has not been turned into waste paper along with them.
02
There is one point that really deserves attention, and it is also the most powerful point in this wave of public condemnation. That is, not all the books purchased by AI companies are ordinary old books that can be bought everywhere. Some of them are indeed rare secondhand books, and even out-of-print books.
This problem cannot be simply glossed over with the statement that "the knowledge remains even if the physical book is gone". Because the value of some books does not only lie in their text. Special editions, unique bindings, author's inscriptions and annotations, and even the collection history of a book and the traces left on it may make the physical object itself have irreplaceable value.
But first of all, we need to make one thing clear: scarcity does not equal preciousness, and being out of print does not equal being a cultural heritage.
If your university teacher self-published 300 copies of *Manual for County-level Sewage Pipe Network Construction* in 1994, there may only be a dozen copies left in the world today, which is of course very scarce and has long been out of print. As for whether it is a treasure of human civilization that must be protected at all costs, the answer is self-evident.
There are a large number of such items in the secondhand book market. The reason why they are hard to find may only be that few copies were printed back then, or they were never reprinted, which does not mean that the loss of each copy represents a certain loss of public culture.
Secondly, even if we narrow the scope to those books whose physical form itself is of great value, the problem is still not that simple.
As long as a book circulates freely as a commodity in the secondhand book market, it is possible to be legally bought by a "crazy person".
This buyer can enshrine the book in a constant temperature and humidity glass cabinet, or use it to prop up a wobbly table. Even more extreme, he can tear off two pages and soak them in milk for breakfast. As long as the book is not protected by special laws, others can only call him a waste of precious things, and no one can stop him.
So why does the situation suddenly change once the buyer is an AI company? It is as if after Anthropic or Amazon pays for the book, although it nominally owns it, it actually has to keep it safe for all humanity.
This kind of debate has already happened among humans themselves.
In 2010, Christie's auctioned a *Book of Hours* produced in France in the 1460s. It was decorated with gold foil, with exquisite calligraphy and 17 full-page illustrations, and was basically intact at the time of auction, finally sold for 25,000 pounds.
A few years later, Elaine Treharne, a scholar of medieval literature at Stanford University, bought the remains of this book. Only 7 of the original 254 pages were left.
She later found out that after the auction at Christie's, the book was actively disassembled by the buyer, a German antique bookseller. The pages were sold one by one, with ordinary pages selling for hundreds of dollars, and beautiful illustrations selling for even higher prices.
But the antique bookseller's subsequent defense was very interesting:
When Christie's held the public auction, museums could come and buy, large collection institutions could come and buy, and wealthy collectors could also come and buy. As everyone saw, no one was willing to pay a higher price in the end, and I paid for it and bought it. Why do you suddenly tell me after the transaction ends that this thing actually belongs to all humanity, and I have to take good care of it for you?
To put it plainly, his words are crude but the logic holds up, and the more you think about it, the more reasonable it becomes.
The same logic applies to AI companies today.
If a certain book is really so precious that slicing even one page is an irreparable cultural loss, then where are the museums? Where are the libraries? Where are the public collection institutions? Where are those humans who care more about the precious assets of humanity?
How come after AI companies actually pay to buy it, everyone suddenly gathers around and says: You are not allowed to destroy it, this belongs to all humanity.
You don't buy it, and you still want to control what I do after I buy it. Isn't that a bit too much?
Of course, this does not mean that it is worthy of praise for AI companies to destroy precious old books. The fact that "it is legally allowed to do so" and our belief that "it is a bad behavior" can coexist completely.
But if a type of book is really important enough that the owner cannot dispose of it at will, what really needs to be changed is probably not the moral standard of AI companies.
03
To make the logic clear, I have been stating the point in a rather extreme way just now.
From the perspective of the common interests of humanity, AI companies' large-scale procurement of secondhand books, especially focusing on niche, out-of-print, and minority language books, is indeed a problem worthy of vigilance.
As mentioned earlier, Anthropic buying millions of books a year is actually nothing in the entire secondhand book market. But if the procurement volume that is a drop in the bucket in the whole market is highly concentrated on rare books, it will still cause large-scale damage in this specific category.
Therefore, the claim that "AI is buying up all secondhand books" is exaggerated, but the concern that "AI companies' procurement may focus on consuming some already fragile book categories" is not unfounded.
But the point is, don't expect AI companies to play the role of saints.
Of course they can do that.
But "it would be best for them to do so" and "society can only count on them to do so" are two completely different things.
Enterprises are enterprises after all. Today a certain company may think it is important to protect rare books and be willing to spend extra money to do non-destructive scanning. Tomorrow, when a new person in charge takes office and finds that this process costs 30% more, it is entirely possible to cut the budget.
Expecting commercial companies to guard cultural heritage for the whole society relying on their conscience in the long run is a bit like asking Sun Wukong to put the headband on himself and recite the tightening spell by himself — of course it is a happy ending if he figures it out occasionally, but system design had better not count on such miracles.
If AI companies' procurement really begins to threaten some books of public value, this is no longer a problem of "how good the corporate morality is", but a public interest issue.
Since it involves public interests, the solution should also come from public rules.
For example, a more complete registration system for rare bibliographies and editions can be established. If a book is confirmed to have very few surviving copies and the physical object itself has clear historical, edition or cultural relic value, it can be included in a certain protection list. AI companies can buy it, but cannot perform destructive scanning. Or after purchasing, they must prioritize non-destructive digitization.
We can also give libraries, museums and public collection institutions a certain priority right. When the commercial procurement system finds a suspected unique copy or a version with very few surviving copies, it will first submit the information to public institutions, giving them a certain period of time to decide whether to collect it. Only when they don't want it can the buyer purchase and process the book.
Secondhand book platforms are not completely free from responsibility. Now that algorithms can accurately find that "there are only three copies of this 1987 book online", the platform can of course trigger an alert when the same book is suddenly purchased centrally by multiple large-volume buyers.
We can even require large-scale purchasers to assume certain information disclosure obligations. If you perform destructive scanning on millions of books a year, you should at least record what editions you have processed, which books are extremely scarce, and contribute this set of data to the public bibliography system.
Which of these solutions is the best is open to discussion, and maybe none of them will be adopted in the end. But at least there should be a set of rules.
04
At this point, there may be one last question left.
When people scold AI companies for "destroying human books" and "devouring human knowledge", what exactly are they scolding? Where does that strong sense of discomfort and fear come from when seeing books having their spines cut off and sent into scanners?
The increasingly tangled feelings humans have towards AI itself cannot be ignored. People are becoming more and more dependent on AI on the one hand, and more and more afraid of it, even more and more resentful towards it on the other hand.
Pew's survey this year shows that the proportion of US adults using ChatGPT has risen from 18% in 2023 to 44% in 2026, and 38% of the employed population has used AI to handle work. But at the same time, 63% of Americans think AI is developing too fast, and 40% expect it to bring negative impacts to society in the future.
This contradiction is particularly obvious among young people.
Gallup surveyed American young people aged 14 to 29 this year, 51% use generative AI at least once a week, and 2