- A bookseller filled an order for 1,000 rare books from a buyer who never had to give a name, and hid an Apple AirTag inside one volume before shipping.
- The tag stopped at an Amazon warehouse near Las Vegas, at a team called VGT3, where bindings are cut off so pages feed through a scanner faster.
- Amazon says it buys books through commercial channels to help develop and improve its products. The scanned text feeds its Nova models.
- A US court has already held that buying a print book and digitising it is fair use, so the seller loses every say the moment the box leaves.
- The part none of the coverage mentions: a thousand units bought by someone who will never buy again poisons the demand data you reorder from.
- Your rights here are thin. Your records are not, and that is where a small seller has something to work with.
The order looked like the best week the shop had all year. One thousand books, paid up front, one shipping address. On Biblio, the marketplace where it was placed, a buyer does not have to say who they are, and this one did not. Something about the list bothered the seller anyway: the titles had nothing in common except that they were old and hard to find. So before the pallet went out, an Apple AirTag went between the pages of one volume.
It stopped moving in Nevada. 404 Media followed the tracker to an Amazon warehouse near Las Vegas and found a team there, VGT3, whose work is to take in printed books and get the text out of them quickly. The fastest way to do that is to remove the spine so the pages can be fed through a scanner loose. The book does not survive that. It is not resold, not donated, not returned. It becomes paper waste, and its contents become AI training data.
For anyone who sells physical goods, this is not really a story about books. It is a story about what a purchase order can be, and about how little you find out.
What did the AirTag actually show?
It showed a route and an endpoint, and the endpoint is what mattered. The shipment travelled from the seller to a large Amazon facility in Las Vegas, where The Decoder reported the receiving team is known internally as VGT3, complete with a logo of a Tyrannosaurus rex holding a book. Workers described the job plainly. One told reporters that scanning books is all they do.
Amazon did not deny the operation. Asked about it, the company said it purchases books through commercial channels to help develop and improve the products and services its customers use, a sentence that is true and answers nothing. Futurism published the same statement alongside the detail that the spines come off before the pages are read.
Nothing here was illegal, and nobody has suggested otherwise. The books were bought at the asking price on an open marketplace. That is the uncomfortable part.
Why is a printed book worth this much to a model?
Because of what it is not. It is not on the web, and it was not written by a machine. Booksellers who have watched these orders come in say the buying looks systematic, worked through by ISBN rather than by subject, which is what you would expect if the goal is coverage rather than reading.
Text printed before 2022 has a property that gets scarcer every month: nothing in it was generated by another model. Anything scraped from the open web now arrives mixed with machine output, and a model trained on its own family's writing gets duller, not sharper. Old paper is one of the few large stocks of text that is guaranteed clean, and it is finite.
Amazon is not vague about the general principle, even if it is vague about the specific pallet. Its own service card for the Nova model family says the Amazon Nova corpus is built from licensed and proprietary data, open source datasets and publicly available data where appropriate, and names the training data corpus as one of only two real levers it has over how the model behaves. When the corpus is one of two levers, buying more corpus is not a side project.
Did the seller have any say after the sale?
No, and this is settled rather than arguable. Section 109 of US copyright law, the first sale doctrine, says that the owner of a lawfully made copy may sell or otherwise dispose of that copy without the copyright owner's authority. Once a specific physical copy is sold, control over that copy goes with it. A bookseller who sells you a first edition cannot tell you not to cut it up, any more than a potter can tell you not to break a bowl.
The scanning question was answered separately, and recently. In Bartz v. Anthropic, a federal judge in California split the case in two. Downloading pirated books to build a permanent library was held to be infringement with no fair use defence. Buying print books and digitising them was held to be fair use, on the reasoning that the company had simply replaced print copies it had paid for with searchable digital ones. The Authors Guild published a sharp response to that half of the ruling, arguing it ignores how publishing actually licenses print, electronic and audio rights separately.
Anthropic ran the same play before Amazon did, under an internal name that surfaced in the litigation, and the pattern was identical: buy print copies on the open market, remove the bindings, scan, discard. The half of that case the company lost was about pirated files rather than purchased paper, and it ended in a settlement of 1.5 billion dollars approved in July 2026, the largest copyright class action settlement of its kind in the United States, a theory the music publishers now suing over Claude's training corpus have borrowed wholesale. Read the two halves together and the incentive they create is obvious. Downloading is now expensive and legally exposed. Buying is cheap, lawful, and leaves the seller with a completed order and no idea what happened next. The market did not correct the behaviour, it just moved it into the purchase ledger, which is where a merchant would see it if a merchant knew to look.
Worth being precise about what the buying is not. It is not a bid for scarcity or for resale value, so none of the usual signals apply. A collector who wants a first edition wants that specific object and will pay for condition. A scanning operation wants the text, which means condition barely matters, the ISBN matters a great deal, and a battered reading copy is as good as a clean one. If you ever wondered why an order ignored the grading you spent time writing, that is the tell. The buyer is not reading your description of the object. They are reading the identifier.
Whatever you think of the reasoning, its practical effect is clear. Buy the physical copy, and scanning it is defensible. Which is exactly the route the Las Vegas pallet took.
Who controls what, once the box has shipped
| Question | Who decides | What the seller can do |
|---|---|---|
| Can the buyer resell the item? | The buyer, under first sale | Nothing after the fact |
| Can the buyer destroy the item? | The buyer, absolutely | Nothing |
| Can the buyer digitise it for internal use? | Courts, and so far the answer leans yes for purchased copies | Nothing, unless a licence was signed at sale |
| Must the buyer identify themselves? | The marketplace's rules | Choose channels that show you the buyer, or sell direct |
| Does the sale count as customer demand? | You do, in your own records | Tag it and exclude it, which is the one lever that pays |
What does a one time bulk order do to your reorder point?
It quietly ruins it, and this is the part of the story that has gone unwritten. Every forecasting method a small shop uses, from a spreadsheet moving average to a hosted demand tool, works on one assumption: that past sales are a sample of future customer interest. A buyer acquiring stock to destroy it is not a sample of anything. They will not reorder, they will not come back next season, and they will not tell you they were there.
Work through what the record now says. That title sold 1,000 units in a week, so it climbs to the top of your bestseller report. Your reorder logic sees a demand spike and raises the safety stock. If you price dynamically, the sell through rate pushes the price up on the units you have left, and genuine customers see a higher number for a book nobody else actually wanted that week. Next quarter you buy in more of a line that has no real audience, and the cash sits on a shelf.
None of that requires a sophisticated system to go wrong. It requires only that the sale be recorded like every other sale, which by default it is. The fix is unglamorous and takes ten minutes: a flag on the order, an exclusion rule in whatever produces your reorder numbers, and a note of the reason. If you want the longer version of why the input data matters more than the model, we wrote about what a demand forecast needs before it is worth trusting, and the same rule holds here. A clean method on contaminated history produces a confident wrong answer.
There is a second order effect worth naming. If you sell on a marketplace whose ranking depends on sales velocity, a bulk purchase moves your listing up for reasons unconnected to shopper interest, then drops it when the artificial demand stops. You inherit both the spike and the cliff.
Your catalogue is text as well
Stock is the visible half. The other half is everything you have published: product descriptions, care instructions, size guides, the long blog post explaining how the thing is made. That material has the same qualities the buyers of old books are paying for. It is specific, it was written by a person who knows the subject, and there is no other copy of it.
The difference is that nobody has to buy it. Web content is collected by crawlers, and the only consent mechanism is a text file most shops have never edited, which is why the argument over whether training on published material counts as fair use matters more to sellers than it first appears. Whether that file changes anything is a fair question, and we looked at the evidence in a piece on whether an llms.txt file does anything measurable. The short answer is that the honest crawlers respect the controls that exist and the rest do not, which makes the file worth writing and not worth relying on.
There is a flip side that cuts the other way. Being readable by machines is now how a shop gets mentioned inside AI answers, which is a real source of buyers rather than a hypothetical one. Blocking everything to protect your copy also removes you from the surfaces where people ask what to buy. Most shops should be selective rather than absolute, and should know which of their pages carry information worth protecting. Our note on how AI content detection reads shop pages now covers where that line tends to fall.
What can you actually do this month?
Less than you would like about your rights, more than you think about your records. Five things are worth the time, and the order matters.
Ask who is receiving a large order. Not as an accusation, as a normal trade question. Where is it going, is it for resale, do they want an invoice in a company name. A buyer with a real business answers in one line, and an anonymous buyer who will not answer at all has already told you something. On the channels where anonymity is the default, this is the only visibility you get.
Price a bulk order at replacement cost, not at list. Sellers of rare goods habitually discount for volume, which is exactly backwards when the units are irreplaceable and the buyer is not price sensitive. If you cannot restock it, the price should reflect that you are giving up the item permanently, not that you are moving inventory.
Flag the order out of your demand history. Whatever produces your reorder numbers, give it a way to ignore an order. If your tooling has no exclusion field, use a separate customer record so you can filter it later. Do this on the day, because in three months you will not remember which week was real.
Write down what you observed, with dates. Quantities, the marketplace, the titles or SKUs, anything the buyer said. If a licensing market for this material ever develops, and the Authors Guild is arguing hard that one should, the sellers with records will be the ones able to make a claim. The sellers with a vague memory of a big order in August will not.
Decide what your published text is for. Some of it exists to be found by machines and quoted back to shoppers. Some of it is genuinely proprietary and should not be handed over for free. Sorting your own pages into those two piles takes an afternoon and it is the only version of this decision you actually control. If you want to see how we treat the same question for the stores built on our platform, our security page sets out how store data is handled and who can reach it.
None of this is a reason to refuse large orders. Bulk buyers are usually exactly what they look like, and turning away a genuine wholesale customer to guard against a rare case is a worse trade than the one you are trying to avoid. The point is to know which kind you are filling, and to keep the sale from lying to you afterwards.
The thing that should worry a seller
Not the destruction of the books, though that is the image everyone will remember. It is the asymmetry. The buyer knew what they were acquiring, why it was valuable, and what they would do with it. The seller knew a price and an address. That gap is the whole story, and it is widening, because the value of specific, human made, hard to find material is rising while the tools for finding out who wants it and why have not changed at all.
The bookseller in this case closed the gap with a twenty nine dollar tracker and some suspicion. That is not a scalable answer, and it should not have to be one. But it is a reminder that the question worth asking about an unusually good order is not whether the money cleared. It is what the buyer thinks they are getting, and whether you agree with their valuation.
If your answer is that you would have charged more had you known, then you already have the information you needed. Charge more next time, and write down why.