~/AI ETHICS/ai-companies-spark-controversy-by-destructively-scanning-rare-books-for-training-data

AI Companies Spark Controversy by Destructively Scanning Rare Books for Training Data

AI companies and digitizers are reportedly using destructive scanning methods—which involve cutting off book bindings to feed pages into high-speed scanners—to acquire training data and digitize rare texts. This practice has ignited intense debates regarding the preservation of physical books, copyright laws, and public access to digitized knowledge. While destructive scanning accelerates the digitization of hard-to-find texts for AI training, it permanently destroys physical copies of potentially irreplaceable books. This highlights a growing tension between the rapid expansion of AI datasets and the preservation of cultural heritage. Destructive scanning involves disbinding a book to run its pages through high-speed sheet-fed scanners, which is faster and cheaper than non-destructive, white-glove archival methods. Critics raise concerns that even rare, out-of-print books under copyright are being destroyed, despite legal rulings that may permit scanning for specific training purposes.

## BACKGROUND

Book scanning can be destructive (cutting the spine to scan loose pages quickly) or non-destructive (using V-shaped cradles and overhead cameras to keep the book intact). The legal landscape of book digitization was heavily shaped by cases like Authors Guild, Inc. v. Google, which ruled that Google's scanning of millions of books for search snippets constituted fair use.

## REFERENCES

## KEYWORDS

#AI Ethics#Copyright Law#Archiving#Digitization

$ subscribe --daily

AI Companies Spark Controversy by Destructively Scanning Rare Books for Training Data | Daily News