AI Companies Reportedly Destroying Physical Books for Training Data
Anna's Archive has raised concerns that AI companies are purchasing large quantities of secondhand physical books and destroying them during high-speed scanning to acquire clean, pre-2022 training data. This practice has sparked debates over digital preservation and the lengths to which AI developers will go to avoid synthetic or machine-generated data. This highlights a growing bottleneck in AI development, where high-quality, human-generated data is becoming scarce, leading companies to physical archives. It also underscores the tension between copyright restrictions, which limit digital sharing, and the preservation of rare physical texts. The scanning process often involves destructive methods, such as cutting the bindings of books to feed pages through high-speed sheet-fed scanners. AI companies target pre-2022 books to ensure the training data is "untouched by machines" and free from AI-generated content.
## BACKGROUND
Destructive book scanning is a digitization method where a book's binding is removed to allow rapid, automated scanning of individual pages, which is faster and cheaper than non-destructive methods that keep the book intact. Anna's Archive is a prominent open-source search engine and metadata index that aggregates content from various shadow libraries like Sci-Hub and Library Genesis to make books and academic papers freely accessible.