Major AI firms are reportedly outsourcing to middlemen to purchase and destroy millions of physical books, raising ethical and cultural concerns.
Islamabad, Pakistan Jul 31, 2026 ALN: To sidestep mounting lawsuits and a severe shortage of clean digital data, major AI companies are secretly purchasing millions of physical secondhand books through anonymous third-party middlemen.
Using industrial hydraulic cutters and high-speed scanners, firms undergo “destructive scanning”—slicing book spines and digitizing the pages before shredding the physical remains to train next-generation large language models (LLMs).
The Scramble for “Clean Data” and the Internet Shortage
As AI models get bigger, tech giants are hitting a major wall: the open internet is running out of high-quality, human-written text. This issue has arisen due to the increasing reliance on data scraped from the internet, which has led to a phenomenon known as model collapse. Model collapse occurs when the quality of the data used to train AI systems deteriorates, primarily because a significant portion of the data becomes AI-generated rather than human-generated. As a result, the models begin to degrade, leading to poorer quality outputs and diminished logical coherence over time.
To combat this issue, AI developers are turning to pre-2022 books, which have become the new gold standard for training data. These physical books, published before the generative AI boom, are considered uncorrupted, expertly edited, and long-form, making them ideal for training more sophisticated AI systems. Unlike the vast amount of content available on the internet, these books provide a reliable source of clean human data that can enhance the performance of AI models.
Rather than licensing digital catalogs directly from major publishers or buying books under their own corporate names, tech firms are employing covert supply networks and third-party logistics firms to buy physical literature in bulk. This strategy allows them to acquire vast quantities of books without drawing attention to their activities.
Key mechanics behind the operation include:
Anonymous Middlemen: Tech companies are hiring third-party vendors and specialized book buyers, such as Canadian buyer Zoom Books, to establish bulk purchasing accounts across used bookstores, library sales, and online platforms like Alibris and Biblio. This approach helps maintain anonymity and allows tech firms to avoid scrutiny regarding their data acquisition practices.
“Project Panama”: Legal documents unsealed in court cases against AI developer Anthropic revealed covert initiatives like “Project Panama.” This project allocated tens of millions of dollars to acquire over two million physical books. Such initiatives underscore the lengths to which AI companies are willing to go to secure high-quality training data.
Hydraulic Slicing and Shredding: Once the books are shipped to temporary warehouses, industrial hydraulic cutting machines are used to slice off the book bindings. The loose pages are then fed through production-level, high-speed scanners to extract digital tokens. After digitization, the remaining paper is sent for pulping and recycling, effectively erasing the physical book from existence.
The “First-Sale” Legal Loophole and Battles
The push toward the destruction of physical books is heavily influenced by recent legal rulings and the evolving dynamics in the United States. AI companies are leveraging an old legal principle known as the first-sale doctrine. Under U.S. law, this doctrine allows individuals to resell, give away, or physically alter a book they legally purchased. By claiming that their actions constitute “transformative use,” tech firms argue that digitizing a book they own for training purposes is protected under Fair Use.
This legal argument also serves as a strategy to avoid piracy lawsuits. The practice of scraping digital “shadow libraries” like LibGen or Books3 has already led to multi-billion-dollar class-action suits from authors and publishers who contend that their intellectual property is being infringed upon. In contrast, by destructively scanning physical copies they own, AI developers believe they have a more legally defensible position for building their training datasets.
Cultural Concerns and Long-Term Feasibility
Despite providing a temporary influx of pristine human text, industry analysts express concerns that destructive scanning is an unsustainable, short-term fix. The implications of this practice extend beyond legal and technical considerations; they also touch upon cultural preservation.
Irreversible Cultural Vandalism: Antiquarian booksellers and archivists have raised alarms about the potential loss of rare, out-of-print, or foreign-language titles. Many of these books represent some of the last surviving physical copies of important works, and their destruction for data extraction purposes is viewed as a form of cultural vandalism. The erasure of these texts not only diminishes the diversity of available literature but also represents a significant loss to cultural heritage.
Finite Supply: With approximately 130 million unique books ever published historically, yielding around 30 to 40 trillion text tokens, the physical literary canon alone cannot indefinitely sustain the exponentially growing data demands of future frontier AI models. As AI systems continue to evolve and require more sophisticated training data, the finite nature of physical books raises questions about the long-term viability of this approach.
Moreover, as the demand for high-quality training data increases, the competition among AI companies is likely to intensify. This could lead to further aggressive tactics in acquiring physical books, potentially exacerbating the cultural concerns surrounding the destruction of literary works.
In conclusion, the practice of destructively scanning physical books for AI training data reflects a complex intersection of legal, ethical, and cultural considerations. While it provides a temporary solution to the challenges posed by data scarcity and model collapse, it raises significant questions about the sustainability of such practices and their implications for cultural preservation. As the AI industry continues to grow, stakeholders must grapple with the balance between technological advancement and the protection of our literary heritage.
To learn more about the latest developments in Software & Platforms, stay updated with our exclusive reports and analyses on AiLensNews.