As artificial intelligence companies compete for increasingly scarce training material, some are reportedly turning to an unexpectedly old-fashioned source: physical books.
AI developers and their contractors are purchasing large quantities of secondhand books, removing their bindings and scanning the pages before destroying or recycling the originals, according to a July 21 report by Gadget Review published through Yahoo News.
The books are not necessarily connected by subject, author or popularity. Orders may contain hundreds or thousands of unrelated titles covering fields such as history, science, law and foreign-language literature. What they have in common is that they provide lengthy, professionally edited material that was generally written before generative AI began flooding the internet with machine-produced text.

Printed books offer uncontaminated human writing
The rapid expansion of generative AI has created a problem for companies seeking fresh data with which to train their models. As more online articles, social-media posts and reference materials are produced or modified by AI, developers face a growing risk of training new systems on the output of earlier systems.
Researchers have warned that repeatedly feeding synthetic material back into AI models can degrade their reliability and reduce the diversity of their responses, a phenomenon commonly described as model collapse.
Success
You are now signed up for our newsletter
Success
Check your email to complete sign up
Printed books, particularly those published before the widespread adoption of generative AI in 2022, offer an alternative. They were written, edited and published through predominantly human processes and often contain information that has never been made freely available online.
BookData.ai, which markets datasets assembled from books, describes published works as “the highest-density source of structured, coherent human thought,” the Gadget Review article noted.
For AI companies, that makes old books valuable not primarily as collectible objects, but as reservoirs of clean, organized language.

Books are processed on an industrial scale
According to the report, companies or intermediaries working under nondisclosure agreements may order books by the pallet. A standard pallet can contain hundreds or even more than 1,000 volumes, while larger purchases can involve tens of thousands of titles.
The books are then prepared for rapid digitization. Their spines are cut away so pages can be fed through high-speed scanners, allowing each volume to be converted into machine-readable text far more quickly than would be possible using nondestructive library scanning equipment.
Once scanned, the physical copies may be pulped or otherwise discarded.
The report also pointed to Anthropic’s previously disclosed book-digitization effort, known as Project Panama. Court filings examined during copyright litigation showed that the company purchased millions of physical books and used a contractor to remove their bindings, scan them and dispose of the paper copies.

Court ruling provided a legal pathway
The practice received significant legal support from a June 2025 ruling in Bartz v. Anthropic.
U.S. District Judge William Alsup concluded that Anthropic’s use of lawfully purchased books to build a searchable digital research library and train its Claude models was transformative under U.S. copyright law. The judge distinguished those purchased books from millions of unauthorized digital copies Anthropic had obtained from pirate libraries.
In describing the company’s scanning process, the court found that converting a legally purchased physical book into a digital copy—while destroying the original and retaining only one replacement copy—could qualify as fair use.
That distinction effectively provided AI developers with a potential legal model: purchase the physical book, digitize it for internal use, eliminate the original and do not distribute the scanned copy as a substitute ebook.
The ruling did not provide blanket permission to acquire copyrighted material illegally. Anthropic continues to face liability over books obtained from unauthorized online collections.
READ MORE:
- OpenAI Says AI Agent Breached Testing Environment, Raising Safety Concerns
- Bessent Warns Chinese AI Firms Could Face US Sanctions Over IP Theft
- AI Data Center Boom Sparks Backlash Across Canada
Destruction raises concerns about rare works
The process becomes more controversial when the purchased material includes scarce, out-of-print or foreign-language titles.
A mass-market novel may exist in thousands of libraries and private collections. A regional history, obscure academic study or decades-old foreign-language publication may survive in only a handful of locations. When one of those copies is purchased and destroyed, it may become harder, or potentially impossible, for members of the public to access the original work.
The Gadget Review report warned that some low-circulation works could disappear from public view after being scanned into proprietary corporate databases.
Other digitization methods demonstrate that destruction is not unavoidable. Libraries have spent decades scanning public-domain works with nondestructive equipment, while Harvard University has worked with technology companies to release a dataset containing nearly one million public-domain books in hundreds of languages.
The growing AI demand for printed books therefore raises a broader question: whether preserving human knowledge should take priority over rapidly converting it into privately controlled training data.
For AI companies, a worn box of old books may represent valuable raw material. For libraries, historians and readers, however, those same volumes may be irreplaceable pieces of the cultural record.