Booksellers suspect AI firms destroy rare books for training data
Ars Technica describes page-tearing scans as cheapest workflow, Internet Archive keeps slow human scanning for fragile volumes
Images
Photo of Ashley Belanger
arstechnica.com
Some booksellers say they suspect AI firms are buying rare books and physically dismantling them to speed up scanning for training data. Ars Technica reports that the process described by sellers involves tearing out pages, cropping them, scanning, and then discarding the paper—an approach presented as the fastest and cheapest way to convert long-form texts into machine-readable datasets. The reporting frames it as part of a broader race among AI developers to secure high-quality writing as training material.
The mechanics matter because they point to a trade-off the market is already making: speed and volume over preservation. Non-destructive scanning exists, but it is slower and more expensive, and it can introduce its own quality problems. Ars Technica notes that Google patented a non-destructive scanning method years ago, yet studies have found issues such as page-curvature distortion and missed pages. At the scale of “millions of titles,” the article says, those frictions become cost items that firms may try to avoid—especially when the books in question are difficult to handle and errors in capture can be tolerated if the goal is statistical training rather than archival fidelity.
The contrast in incentives is clearest in the Internet Archive’s scanning operation, which Ars Technica revisits through earlier documentation of its workflow. The Archive treats scanning as careful manual labour, emphasising slow handling, repeated checks, and systems that stop the process when a page is skipped or an image is blurry. Staff described using clean, dry hands to turn fragile pages and procedures to catch fold-outs and inserts, aiming for what they describe as zero errors. The article notes that the Archive tested automated commercial scanners, including vacuum-powered page-turning arms, and found they performed poorly on brittle or rare volumes.
If AI firms are indeed dismantling books, the costs are not borne by the scanning operation alone. Rare-book markets depend on provenance, condition, and scarcity; removing physical copies to create private datasets shifts value from an object that can be resold, loaned, or donated into a digital asset that can be used indefinitely inside a model. Ars Technica reports that book lovers and sellers worry the practice is happening at a larger scale than the limited anecdotes suggest, but the article does not identify specific companies conducting the destruction.
The cheapest workflow described is also the least reversible. Once the pages are torn out and discarded, the “dataset” remains, and the book does not.
Ars Technica’s account ends with a detail from the Archive’s scanning floor: a foot pedal lifting scanner glass, one page at a time.