Why AI book scanning is putting rare volumes under scrutiny

AI book scanning has drawn alarm because the fastest method can mean cutting bindings, scanning pages, and discarding the physical copy. The Internet Archive offers a different model: slow, human-led digitization designed to protect fragile and rare books.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 3 ►

The story centers on AI demand turning books into disposable training inputs and risking cultural preservation, rather than autonomous AI danger.

Why AI book scanning is putting rare volumes under scrutiny

AI companies want books because books contain long-form, high-quality writing that can help train models. The conflict is over how those books are turned into data: quickly and destructively, or slowly enough to preserve the physical copy.

That question has become sharper as booksellers report unusual bulk orders and book lovers worry that rare or fragile works could be treated as disposable input for AI systems.

The speed problem behind AI book scanning

The fastest route from printed page to training data is also the one that alarms preservation-minded readers. A book can be stripped from its binding, its pages prepared for scanning, and the physical object discarded after the images are captured.

That approach is cheap and fast, which matters to AI firms racing to improve models with large collections of books. But it also creates a hard preservation risk: once a physical copy is destroyed, it cannot serve readers, collectors, libraries, or historians in the same way again.

The concern is not only theoretical. A lawsuit in summer 2025 revealed that Anthropic had destroyed millions of print books to train its AI models. The article notes that there is no indication Anthropic destroyed rare books, and that the company has denied doing so in statements.

Still, the broader practice has changed how book buyers are being watched. Rare booksellers are increasingly alert to orders that look less like collecting and more like sourcing inventory for mass scanning.

A slower method already exists

Destructive scanning is not the only way to digitize a book. Google patented a non-destructive book-scanning technology in 2009, showing that faster digital capture does not have to require cutting the spine from a volume.

That technology has limits. The curve of a page can distort text, pages may be missed, and mistakes can appear when workers move too fast. Wired reported that glitches have included hands blocking pages during scanning.

Those trade-offs help explain why preservation work is not only a technical task. The Internet Archive, which works with libraries to preserve aging collections, treats scanning old books as careful human labor. The goal is to reduce handling while still creating usable digital copies.

That method may be too slow or costly for companies seeking the cheapest way to scan millions of titles. But for brittle books, rare volumes, and special collections, speed is not the only metric that matters.

How the Internet Archive protects fragile books

An Internet Archive post from 2021 describes a process built around patience. Eliza Zhang, a book scanner who has been with the Archive since 2010, scans books one page at a time and checks the quality as she goes.

The Archive had tried automation, including “commercial book scanners that feature a vacuum-powered page-turning arm.” But the organization said those systems did not work well for the kinds of fragile and rare materials its library partners ask it to digitize.

“clean, dry human hands are the best way to turn pages.”

That explanation came from Andrea Mills, who helps lead the Archive’s book-scanning operations. Zhang described the job as one that demands “keen concentration,” especially because pages in “very old, fragile books” can be “paper thin.”

The process is deliberate. Zhang raises the scanner glass with a foot pedal, turns a page, adjusts the cameras, and checks that the page can be read. When a book includes fold-outs, she marks them so they can be scanned separately and not lost while the main pages are documented.

The Archive’s proprietary software also checks the work. If a page is skipped or an image is too blurry, the process stops and asks for another scan. The point is to avoid putting a rare book through the process again.

At the time of the 2021 post, Zhang had scanned “more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books),” with the goal of “zero errors.” Chris Freeland, the Internet Archive’s director of library services, told Ars that the post remains “still the best description of our scanning process today.”

Booksellers are watching bulk orders

Reports about AI and books have also put rare booksellers on alert. The Atlantic reported that social media “raged” after two recent reports suggested AI could already be putting rare books at risk.

One Telegraph report accused Silicon Valley of destroying millions of rare books and “shredding the originals.” Separately, 404 Media reported that ISBNdb, a book-database company, was advertising that it could help AI firms source books in bulk.

The backlash was not unexpected. ISBNdb reportedly warned clients that the optics were bad. Anthropic also used the codename “Project Panama” for destructive book scanning in an effort to keep the work hidden from the public.

Booksellers now have reasons to question unusual purchases. The Irish Times reported that an Irish bookstore called Kennys flagged a “bananas” order for 5,000 obscure titles. The order stood out because the buyer did not try to haggle on price, which the Times described as an obvious “red flag” that AI might be involved.

Some sellers may benefit from large purchases in the short term. But the Times reported that booksellers fear selling to AI firms that destroy collections indiscriminately “may be signing their own death warrant.”

The unresolved trust problem

Not every AI-related book project is destructive. OpenAI and Microsoft are working with Harvard librarians on an effort to train AI models on about 1 million public-domain books dating back to the 15th century.

Elon Musk’s xAI has also publicly said it will not destroy rare books to train AI. In a post on X, Musk said he “asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning.”

Yet critics remain skeptical, partly because the word “rare” can leave room for interpretation. Some have noted that Musk did not say xAI would avoid destroying any books, only rare ones.

That is the core tension now facing AI book scanning. Companies want the knowledge contained in books, but the way they acquire it can either preserve cultural objects or erase them from the physical world. The Internet Archive’s process shows that careful digitization is possible. The unanswered question is whether AI firms will accept the slower path when speed and scale are the prize.