September 2, 2026
The Warehouse Where Books Get Cut Apart for AI
A tracker hidden in a box of rare books led to an Amazon building in Las Vegas where the spines come off and the pages go in a bin. Used bookstores from Toronto to Vancouver are getting strange orders. And a lawsuit says the most valuable thing about an old book is that a machine didn't write it.
The facts:
- 404 Media put a tracking device in a shipment of rare books a dealer suspected was headed to an AI company; it landed at an Amazon warehouse in Las Vegas where, an employee says, workers slice the spines off with a guarded blade, feed the loose pages through 20 to 25 scanners "like machines that count cash," and throw the pages into six-foot cardboard bins (404 Media)
- The books came from everywhere: sealed pallets from Japan, German and Russian titles, liquidated library stock, University of London material, and stapled papers marked as presented to Parliament; staff were first told it was "for Kindles" (404 Media)
- Anthropic's internal effort, called Project Panama, described itself as a plan to "destructively scan all the books in the world" (Fortune)
- A federal judge ruled last year that training on books is fair use and that scanning a bought copy and destroying it is too, but that downloading pirated copies is not; Anthropic paid $1.5 billion to settle with authors over 7 million pirated books (Reason, Variety)
- On Friday, Sony Music Publishing and Warner Chappell sued Anthropic and named its CEO and a co-founder personally, citing internal chats praising a pirate library and arguing $1.5 billion "is obviously not a large enough settlement" for a company valued at $2 trillion; they want up to $150,000 per song, including "Hallelujah," "Uptown Funk" and "I Am the Walrus" (Ars Technica, Fortune, Al Jazeera)
- Anthropic's reply: "the third lawsuit from the same lawyers, recycling allegations," and training is fair use (Variety)
- Canadian used bookstores report bulk orders since June for random lists of old technical, medical and math books, priced $5 to $30, shipped to a logistics warehouse in Illinois; one Toronto owner now refuses all bulk orders: "there's something not right here" (CBC)
- Books printed before 2022 are worth more to AI companies because they're guaranteed free of AI-written text, which degrades a model when it trains on its own output (CBC)
- Reason's counterargument: the books being bought have ISBNs, so nothing older than 1970, nothing antiquarian, and a one-to-one scan preserves the text; the FTC has been asked to investigate anyway (Reason)
- Jason Isbell and other musicians sued the music generator Suno on Monday, not over copyright but over identity, claiming the model holds "voiceprints" users can reach by asking for an artist's "tone and phrasing" (Variety)
- A student scraped 12 million images from Cara, an art site built to keep AI out, for under $10, bragged on Reddit, then apologized and is now helping the site's founder build protections (Wired)
Picture the room. Twenty-some scanners riffling pages at the speed of a bank counting bills. A person at each cutting station sliding a book under a blade, pulling their fingers back, pressing the button. Behind them a cardboard box taller than a man, filling up with loose paper that used to be the University of London's library.
The employee who described it to 404 Media said the thing that bothered her wasn't the machines. It was that some of the books looked rare and she wished she could take them home.
why the books
A model learns from text, and the text it can't have is the text nobody posted online. Old textbooks, out-of-print technical manuals, sheet music, parliamentary papers. Every book is also a certificate: if it was printed before 2022, no machine wrote it, and a model that trains on machine writing gets worse. So the demand isn't for classics. It's for anything with a copyright date and a spine.
That's what the used bookstores in Toronto and Vancouver were seeing without knowing it. A dozen sales a day of dated nonfiction that hadn't moved in years, at real money. A LinkedIn message from a data company "exploring sourcing partners." A buyer in British Columbia whose new growth officer came from a UK wholesaler named in Anthropic's court records. The Toronto owner's answer is now no unless he knows you. "We are booksellers. Books shouldn't be destroyed."
Reason's rebuttal is fair and worth stating plainly: these are ISBN-era books, mostly with plenty of copies around, and a scan preserves the words better than a damp basement does. Nobody is feeding a Shakespeare folio into a cutter. The FTC letter calling this anticompetitive is a stretch.
Both things can be true. The text is preserved, and the bookseller is being used as a supply chain by a customer who won't say who he is.
the money part
The courts have settled the shape of this, at least for now. Buy the book, cut it, scan it, train on it: legal. Download seven million pirated copies: not legal, and that cost Anthropic $1.5 billion. Friday's lawsuit from Sony and Warner is the music version of the same fight, with the same lawyers, and it goes after the CEO by name. The publishers' argument is simply that a billion and a half is a rounding error at a $2 trillion valuation, so the fine didn't work.
The Suno case is the more interesting one, because it isn't about copyright at all. Isbell's lawyers say the model holds something like a voiceprint, and you can get at it by asking for his "tone and phrasing" instead of his name. A musician can sell the rights to a recording and still own being that musician. Whether that survives a generator is a question no court has answered.
the student
The Wired story is the small human one. A student pulled every public image off Cara, an art site whose whole reason to exist is keeping AI away, for less than ten dollars, and posted about it on Reddit to make artists angry. It worked. People deleted their portfolios. Then he read what they wrote, felt sick about it, deleted the dataset, and is now in the site's Discord helping fix the holes.
He also said the quiet part: twelve million images is nothing, the big labs don't need it, and the sites that get scraped hardest are the big ones where the artists had come from.
The employee in Las Vegas said the process changed every day and nobody seemed to know what the books were for. The company's spokesperson called the lawsuit recycled. The bookstore owner in Toronto trained his whole staff to say no. The pages are already in the bin.
Sources: 404 Media, CBC, Fortune, Reason, Variety, Ars Technica, Al Jazeera, The Decoder, Wired.