AI firms are buying up old books because they are the last slop-free data left
What happened
AI firms are buying up millions of old physical books to use as clean training data. Companies like ISBNdb pitch these old books as a final source of slop-free text for AI training, meaning data that isn’t polluted by internet spam, misinformation, or copyright noise. The process involves physically slicing the spines off books to digitize the contents. The data brokers and labs involved keep buyers anonymous.
Why it matters
Most modern AI training datasets scrape the web, which introduces vast amounts of junk data and copyright complications. Old printed books offer high-quality, vetted content that can train language models with fewer errors and cleaner semantic grounding. This approach pressures library inventories and raises ethical questions about destroying physical books at scale. It also signals AI companies are willing to pay a premium for pristine, offline data sources as web data grows less reliable and more legally fraught.
What to watch next
Watch for potential legal scrutiny or pushback from libraries, authors, and collectors as this practice scales up. The black-box nature of which buyers fund these conversions will likely fuel further debates on transparency and data provenance in AI training. The cost of acquiring and digitizing physical books may also shift AI economics, favoring players with deep pockets. Finally, expect alternative “clean” datasets to compete with book dumps or new licensing models for digital archives.
AI Quick Briefs Editorial Desk