Content
Recent Posts
Why AI Companies Are Tearing Through Millions of Books

Newly unsealed court documents have revealed that Anthropic, the company behind Claude, quietly bought millions of used books, scanned every page, and destroyed the physical copies to build a massive library of training data. The initiative, known internally as Project Panama, offers one of the clearest looks yet at how AI companies are sourcing the human-written content needed to train increasingly advanced AI models, while raising new questions about the legal and ethical limits of that process.
According to the court records, Anthropic purchased millions of printed books from suppliers, including Better World Books and World of Books. Contractors removed the bindings, scanned the loose pages using high-speed industrial scanners, converted the pages into searchable digital text using optical character recognition (OCR), and recycled the paper. The resulting digital library was used to help train the company's AI models, including those behind Claude.
The project quickly drew criticism from authors and readers, not because the books were stolen, but because they were intentionally destroyed after being digitized. While destructive scanning has long been used by libraries and digitization projects, Project Panama has raised broader questions about how far AI companies should go in their race to collect high-quality training data.
One of the most detailed accounts of the project came from The Washington Post's investigation, which reviewed thousands of pages of internal documents and court records.
Content
Why the Internet Is No Longer Enough
At first glance, buying millions of physical books might seem unnecessary. AI companies already have access to enormous amounts of information online, including websites, forums, research papers, and news articles. The problem is that today's internet is very different from the one that helped train the first generation of large language models.
Over the past few years, AI-generated content has spread rapidly across the web. Blog posts, product descriptions, social media posts, and even self-published books are increasingly being created with AI. Researchers have warned that repeatedly training AI models on AI-generated material could reduce the quality of future models, a problem often referred to as model collapse.
Printed books offer something that is becoming harder to find online: professionally edited, long-form writing created by humans before the rise of generative AI. Internal documents discussed during the Anthropic case showed employees viewed books as an important source of high-quality writing that could help teach AI models how to write more naturally.
That helps explain why physical books have become valuable again. Their worth is no longer limited to being read. They have also become a source of training data for increasingly competitive AI systems.
The Legal Questions & Ethical Debate
The court's ruling highlighted an important distinction between owning a book and copying its contents. In June 2025, U.S. District Judge William Alsup ruled that Anthropic's use of legally purchased print books for AI training qualified as fair use under U.S. copyright law because the process was considered transformative. At the same time, the judge treated Anthropic's separate collection of millions of pirated digital books as a different matter. Claims related to those pirated copies were later resolved through a settlement, while the ruling on legally purchased books remains one of the most closely watched decisions in AI copyright law.
Although the legal picture has become clearer, the ethical debate is far from settled. Many authors argue that purchasing a single used copy of a book should not allow a company to use that work to train a commercial AI system without providing additional compensation to the person who created it. Others point to the destruction of the books themselves. While the words are preserved digitally, the physical copies are permanently lost once they pass through the scanners.
That distinction has become one of the biggest criticisms of Project Panama. Traditional digitization projects often focused on preserving books or making them more accessible. Anthropic's project had a different goal. The books were treated primarily as a way to extract information for AI training rather than as physical works worth preserving.
Court records also showed that Anthropic wanted to keep Project Panama out of the public spotlight, a detail that has added to public scrutiny. Critics argue that this suggests the company anticipated a negative reaction, even if it believed its approach would ultimately withstand legal challenges.
A Wider Challenge for the AI Industry
Anthropic is not the only AI company facing questions about how it gathers training data. OpenAI, Meta, Google, Microsoft, and several other AI developers continue to face lawsuits from authors, publishers, artists, and media organizations over the use of copyrighted works in AI systems.
While each case involves different legal arguments, they all point to the same challenge. Building better AI models requires enormous amounts of high-quality human-created content, and obtaining that content is becoming increasingly difficult.
Project Panama illustrates how quickly books have taken on a new role in the AI era. Instead of serving only as something to read, collect, or preserve, they are becoming valuable sources of data for the next generation of artificial intelligence. Whether that shift is simply a natural consequence of technological progress or a sign that copyright laws need to evolve is a debate that is likely to continue long after the Project Panama documents fade from the headlines.
For more articles like this, visit our Tech News page!