Skip to main content

Amazon Is Literally Slicing Books Apart So AI Can Read Them. Yes, Really

Amazon's book-destruction operation and the AI industry's data problem

Published: August 22, 2026

In August 2026, investigative reporter Emanuel Maiberg of 404 Media planted a $29 Apple AirTag inside a shipment of approximately 1,000 rare books and tracked it as it traveled through Milwaukee and Kenosha, Wisconsin, across Colorado, and finally to the north end of an Amazon warehouse complex in Las Vegas designated LAS8. There, in a unit Amazon internally calls VGT3, employees cut the spines off books, feed the loose pages through high-speed scanners, and discard the physical copies. According to Amazon worker accounts, the scanned text is used to develop and train AI models, though Amazon's official statement does not confirm this directly.

The investigation documented at least one large-scale instance of what booksellers had suspected for months: that Amazon is purchasing rare and out-of-print books in bulk, digitizing them through destructive scanning, and discarding the originals. But the story goes beyond one company's practices. It points to a problem in the AI industry's relationship with human knowledge that existing copyright law was not designed to address.

What the investigation actually established

The 404 Media investigation, published August 17, 2026, relied on a cooperative rare-book seller on Biblio, an independent marketplace where buyer identities remain anonymous. The seller had received a bulk order for roughly 1,000 titles and suspected it was destined for an AI training pipeline. Maiberg placed an AirTag in one volume and tracked its movement over several weeks.

The tagged book's journey was methodical. It first flew to Milwaukee, sat for two weeks in a distribution warehouse outside Kenosha, then traveled by truck westward with a stop in Grand Junction, Colorado, before arriving at LAS8 in northeast Las Vegas. LAS8 is primarily a print-on-demand facility, a place that creates books. But its north end houses a separate operation.

Amazon employees posting on an internal workers' forum identified the unit by its code name: VGT3. The entrance has a logo of a Tyrannosaurus rex holding an open book, an image that, as Futurism noted, looks remarkably like AI-generated art. Worker accounts describe a focused operation: pallets of books arrive, staff slice the spines to separate pages for high-speed scanning, and the physical copies are discarded once digitized.

One employee wrote on the forum: "All we do is scan books." Another described the work as "nice" with initially relaxed pace before production rate targets were introduced.

Amazon's response to the reporting was notably narrow. The company told both 404 Media and Ars Technica:

"Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."
Amazon, statement to 404 Media and Ars Technica

The statement does not mention AI, does not deny destructive scanning, and does not address the fate of the physical books. According to TechTimes reporting, Amazon had previously denied engaging in destructive scanning, a denial the AirTag evidence appears to contradict, though Amazon has not publicly addressed the discrepancy.

The ISBN detail

One operational detail carries more analytical weight than the scanning itself. According to worker accounts reviewed by 404 Media, VGT3 staff scan each book's ISBN barcode before scanning its content. The International Standard Book Number, a 13-digit identifier standardized as ISO 2108 and covering more than 150 countries, is the global cataloging system that makes every commercially published book uniquely identifiable.

This sequencing matters because it suggests Amazon is building a record of what it has processed. In principle, such records could help an organization identify which ISBNs it has already acquired and which remain absent from its collection, allowing systematic, catalog-driven purchasing rather than random acquisition.

Booksellers across the United States and Europe have described exactly this pattern: bulk orders with no coherent subject matter, unified only by ISBNs, placed by anonymous buyers indifferent to condition, edition, or collectible value. Fortune reported the case of Dutch antiquarian bookseller Pieter de Vries, who received a request for more than 3,000 books from a company calling itself "2077AI", a spreadsheet listing titles mostly published between 2020 and 2021 by academic publishers like Elsevier, Wiley, and Oxford University Press. De Vries dismissed it as spam. Several other Dutch booksellers received the same request and similarly disregarded it.

Why physical books?

The obvious question is why Amazon would destroy physical books when digital text is abundant. The answer involves three factors.

Using pirated online book archives, known as "shadow libraries," has exposed AI companies to legal risk. The case Bartz v. Anthropic PBC illustrates this clearly. Court documents revealed that Anthropic had downloaded over 7 million books from pirate sites Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi). In June 2025, Judge William Alsup of the Northern District of California ruled that Anthropic's use of legally acquired books for AI training was "quintessentially transformative" and protected as fair use, but simultaneously held that downloading and keeping pirated copies was not fair use. Anthropic eventually settled for $1.5 billion, covering 482,460 works, in what became the largest copyright settlement in U.S. history.

FactorWhy it drives physical-book acquisition
Shadow-library liabilityPirated downloads expose AI companies to massive copyright lawsuits. Legally purchased physical books are a more defensible training source.
AI saturation of the internetTraining on AI-generated text can cause "model collapse." Pre-2022 books are less likely to contain AI-generated content.
Never-digitized reservoirGoogle estimated ~130 million unique book titles exist; only a fraction have been digitized. The rest exist only in physical form.

Buying physical books through legitimate commercial channels reduces legal liability. The first-sale doctrine generally allows an owner to dispose of a particular physical copy, including by destroying it. It does not, by itself, resolve whether creating a digital reproduction of the copyrighted work for commercial AI training is lawful. That question has not been tested at the appellate level, but the current legal environment treats legally purchased physical books as a more defensible training source than pirated downloads.

Scale and economics

The scale of the operation is difficult to establish precisely, because neither Amazon nor other AI companies have disclosed their book-acquisition volumes. However, several data points suggest the operation is substantial.

Court documents from the Anthropic case revealed that Anthropic planned to scan between 500,000 and 2 million books over a single six-month vendor contract. An internal Anthropic planning document stated:

"Project Panama is our effort to destructively scan all the books in the world. We don't want it to be known that we are working on this."
Internal Anthropic planning document, cited in court filings

Amazon's operation has not been linked to Anthropic, but the internal document shows the scale of ambition within the AI industry. Booksellers report anonymous orders ranging from hundreds to thousands of titles per transaction. The phenomenon has driven a documented increase in sales of physical books printed before 2022, as AI companies race to acquire them.

The economics, from the AI companies' perspective, are favorable, though the exact figures are not publicly known. Used books in bulk can be purchased for relatively low per-unit costs. Unlike licensing digital content from publishers, buying physical copies from the secondary market requires no negotiations, no royalty payments, and no ongoing relationships with rights holders.

The preservation question

The backlash from historians, archivists, and book lovers centers on a straightforward concern: once a physical book is sliced apart and scanned, that specific artifact, with its unique provenance, marginalia, binding, and material history, is permanently destroyed.

The OECD flagged Amazon's practice as raising "copyright and heritage concerns." The Guardian's Kathryn James wrote:

"The risk with generative AI is that we cede the means of production of our large language lives: that we turn from creators to consumers."
Kathryn James, The Guardian

Historians have drawn parallels to other episodes of deliberate book destruction, though the comparison is imperfect: Amazon's motivation is commercial, not ideological. The preservation concern is particularly sharp for rare and out-of-print titles. Some such works may exist in only a small number of surviving copies. Each copy destroyed reduces the total supply permanently, and unlike a digital file, a destroyed physical book cannot be recovered.

The unresolved legal question

The Bartz v. Anthropic ruling established that AI training on legally purchased books constitutes fair use. But the ruling left several questions open that are directly relevant to Amazon's operation.

First, the ruling addressed training, the process of feeding text into a model to adjust its parameters. It did not specifically address whether the act of creating a permanent digital copy of a copyrighted book, solely for the purpose of building a commercial dataset, constitutes a separate act of reproduction that might require its own fair-use analysis.

Second, the ruling applied to the specific facts of Anthropic's purchasing practices. Whether a systematic, industrial-scale operation designed to scan every published book, essentially creating a comprehensive digital library, would receive the same legal treatment is an open question.

Third, the destruction of the physical copies, while legally permitted under the first-sale doctrine, creates a practical asymmetry: the company gains a permanent digital asset while the world loses a physical artifact. Whether courts will consider this asymmetry in future rulings remains to be seen.

What this reveals about the AI industry

Amazon's VGT3 operation is not an isolated case. It is one example of a broader scramble for training data that has outpaced both the legal framework and the cultural norms around knowledge preservation.

Large language models require vast quantities of high-quality text. The readily accessible internet has already been extensively used for AI training. Physical books, particularly those never digitized, are a remaining reservoir of human-authored, professionally edited text. The incentive to acquire and digitize that reservoir is large. The legal tools to prevent it are, at present, limited.

As AI models become more capable and more commercially valuable, the pressure to acquire training data will only increase. Whether the response comes through new legislation, revised copyright doctrine, industry self-regulation, or some combination remains an open question.

Comments

Popular posts from this blog

Fire insurance claim in america

Fire Insurance Claim in America: Complete Step-by-Step Process Filing a fire insurance claim can feel overwhelming, especially when you're dealing with the aftermath of a fire. However, knowing the right steps to take and what to expect can make the process much smoother and help you get the settlement you deserve. This guide walks you through the entire fire insurance claims process, from the moment the fire is extinguished to receiving your final payment. 1. Immediate Actions: The First 24-48 Hours Your safety and the safety of others is the top priority. Do not re-enter the property until the fire department has declared it safe – entering prematurely could expose you to structural dangers or toxic fumes. Once the scene is secure: 📞 Notify Your Insurance Company Immediately. Contact your insurance provider as soon as possible, ideally within 24 to 48 hours of the incident. Have your policy number ready and provide a clear, factual account of the incident, including the...

America's Top Stories August 2026: Iran Deal, VOA Shakeup & the Week That Shook Washington

August 2026 opened with a lot happening at once. Foreign policy, domestic political fights, economic jitters, and a set of primaries that could reshape the Democratic Party — it's all colliding at the same time, and honestly, it's a lot to keep track of. Iran and the Strait of Hormuz: From War to "Imminent" Deal The biggest story right now is the U.S.-Iran conflict, and specifically the wild swing from escalation to diplomacy over the weekend. Back on February 28, President Trump announced what he called "major combat operations" against Iran. Joint U.S.-Israeli strikes hit military, government, and infrastructure sites across the country. The fight has centered on the Strait of Hormuz — that chokepoint through which a huge share of the world's oil passes. Then, on August 3, Trump did something unexpected. Speaking from the Oval Office, he said he was halting planned attacks because "the perimeters of a deal" had been reached. He call...

Powerball Jackpot Hits $1 Billion: How a Single Illinois Ticket Won the Biggest Prize of 2026 — and What Happens Next

Powerball Jackpot Hits $1 Billion: How a Single Illinois Ticket Won the Biggest Prize of 2026 — and What Happens Next   Somewhere in Illinois, there is a slip of paper worth $1.040 billion. Nobody has come forward to claim it yet, and the person holding it may not even know it yet. But after 43 consecutive drawings without a jackpot winner, the longest drought of the year finally ended Wednesday night when a single ticket sold in the state matched every number drawn. The prize is the biggest Powerball jackpot of 2026, the eighth-largest in the history of the game, and the 15th U.S. lottery jackpot across Powerball and Mega Millions combined to cross the billion-dollar line. It also nudged past the $980 million Mega Millions prize won by a ticket sold in Georgia last fall. For the tens of millions of people who bought tickets as the pot swelled, the anticipation is over. For the one person holding the winning numbers, an entirely different set of questions is just beginning ...