Doctorcrypto About RSS Subscribe
Doctorcrypto
HomeGuides › AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots
Guides

AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots

By Malik Sokolov · · 2 min read

Major artificial intelligence developers are quietly acquiring physical books by the millions, cutting them from their bindings, scanning every page to build AI training datasets, and then throwing the originals away, according to reports on how tech companies are sourcing text for their language models.

How the Practice Works

The process, sometimes dubbed "destructive scanning," involves buying large quantities of used and new books, slicing the spines off so pages lie flat, and running them through high-speed scanners. Once digitized, the physical copies are discarded. The resulting text becomes fuel for the large language models that power chatbots and other generative AI systems.

The approach offers AI companies a straightforward way to obtain clean, high-quality written material. Books represent professionally edited, coherent long-form text — exactly the kind of data that helps train models to produce fluent, well-structured responses.

Millions of physical books are being sliced apart and thrown out to teach machines how to write.

Why Companies Are Doing It

The scramble for training data has intensified as AI firms compete to build ever-larger and more capable models. Publicly available web text has already been heavily mined, and much of it is low quality or legally fraught. Books provide a richer alternative, and buying physical copies can sidestep some of the licensing complications tied to digital sources.

Legal considerations play a significant role. Rather than negotiating rights for e-books or scraping copyrighted digital text, some companies view purchasing and scanning physical books as a cleaner path — though the copyright implications of using scanned book content for training remain contested and are the subject of ongoing lawsuits.

Key factors driving the trend include:

  • A growing hunger for high-quality, professionally edited text
  • Concerns over the legal risks of scraping digital content
  • The relative ease of acquiring physical books in bulk

The Broader Debate

Critics have drawn uncomfortable comparisons to book burning, arguing that destroying physical volumes to feed machines carries troubling symbolism even if the text survives in digital form. Authors and publishers have raised concerns about whether their work is being used without consent or compensation.

The controversy sits within a wider fight over how AI systems are trained and who benefits. As lawsuits and regulatory scrutiny mount, the question of how companies obtain — and treat — the written material behind their models is likely to remain a flashpoint in the debate over the future of artificial intelligence.

Was this useful?👍 Yes👎 No