BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Industry/Washington Told a Court That AI Training Is Fair U…
IndustrySeptember 3, 2026
Read · 5 min
ai training copyright · fair use

Washington Told a Court That AI Training Is Fair Use

The US government backed fair use for AI training. What a small seller controls, what changed this week, and what did not.

Key takeaways
  • On 2 September 2026 the US Justice Department filed a statement of interest backing the argument that training a language model on copyrighted text is fair use. It is advocacy in a live case, not a ruling.
  • The filing sits in the consolidated New York litigation that includes The New York Times, Ziff Davis and a separate action brought by 400 local newspapers.
  • The same federal government reached a colder conclusion fifteen months earlier. The Copyright Office rejected the idea that training automatically counts as a new and different use.
  • Nothing about your obligations as a seller changed this week. What changed is the direction of the argument you will be reading about for the next two years.
  • The controls you actually hold are small and worth setting anyway: crawler directives on a site you own, and the terms of any marketplace where you do not.
  • Blocking Google's training crawler does not cost you search visibility. Google documents that plainly.

You photograph your products. You write the descriptions, the shipping page, the answers to the fifteen questions customers keep asking. That material sits on a public URL because it has to, and it is exactly the kind of text and image a model gets trained on. Whether that is legal has been argued in court for nearly three years. This week the United States government took a side.

On 2 September 2026 the Justice Department filed a statement of interest in the Southern District of New York, supporting OpenAI's position that copying written works to train a large language model can qualify as fair use. The filing landed in the consolidated publisher litigation that includes The New York Times, which sued in December 2023, Ziff Davis, and a separate action brought on behalf of 400 local newspapers.

A statement of interest is not a judgment. It is the executive branch telling a judge how it would prefer the law to be read. It carries weight because of who signed it and none of the force of a decision. Read what follows with that distinction held firmly, because most of the coverage will collapse it.

What did the Justice Department actually file?

It filed a brief arguing that the act of training should be judged separately from how the training data was obtained and separately again from what the model later produces. That separation is the whole engine of the argument, and it is the part worth understanding even if you never read another line about this case.

Two other threads run through it. The department argued that rules making it harder to build a strong American AI industry would threaten national security and hand an advantage to foreign adversaries operating under looser constraints. It also argued that a licensing requirement would price training out of reach for everyone except the largest technology companies, concentrating the field rather than opening it.

The New York Times rejected the reasoning. Its spokesperson Graham James framed the position as the administration siding with trillion dollar companies over the people who produce the work. That is the fight in one sentence, and no filing settles it.

Diagram comparing the Justice Department September 2026 position on AI training with the Copyright Office May 2025 analysis of the same fair use question
Two arms of the same federal government, reading the same question differently, fifteen months apart.

Does anything change for your shop this month?

No. Your obligations are identical to what they were in August, and so are your options. A brief filed by one party's supporter does not create a rule, does not settle a case, and does not give anyone permission to do anything they could not do before.

What it does change is the weather. Judges read these filings. So do the lawyers advising the platforms you sell on, and the vendors whose terms you accept when you switch on a new tool. If the reasoning holds, the practical answer to "can my published catalogue be used to train a model" drifts further toward yes, and the remedy drifts further away from copyright law and toward the technical and contractual controls you set yourself.

Three questions that keep getting merged into one

Almost every argument about this subject is confusing because it treats one question as three. Pulling them apart is the single most useful thing a non lawyer can do here, because your exposure is different in each row.

The questionWhat the September filing arguesWhat stays contestedWhy it matters to a seller
How the material was obtainedJudged on its own facts, not excused by the training that followsWhether a given dataset was assembled lawfullyScraping your site is a different act from training on the result
The training itselfGenerally fair use, because the model learns statistical structure rather than republishing worksWhether learning at this scale counts in law as a new and different useThis is the part your robots directives speak to, before it happens
What the model outputsA separate question, assessed on the outputMemorisation, near copies, and who is liable for themIf a model emits your copy, that is an output claim, not a training claim
Whether a licence market existsMandatory licensing would concentrate power in the largest labsWhether feasible licensing should count against fair useDecides if small publishers ever get paid, or only large ones

Keep that table in mind next time a headline says a court approved AI training. Usually one row moved.

Why the Copyright Office reached a different answer

Here is the detail most coverage leaves out. The same federal government has already published a careful analysis of this question, and it did not land where the Justice Department landed.

On 9 May 2025 the Copyright Office released the third part of its report on copyright and artificial intelligence, covering generative AI training. Its conclusions ran against the industry's favourite framing. The office rejected the claim that training automatically counts as a new and different use, warning that when a system produces expressive content competing with the works it learned from, that label gets harder to earn. It flagged the risk that the speed and volume of generated material dilutes the market for work of the same kind. And it said that where licensing is feasible, unlicensed use weighs against fair use rather than for it.

These two documents do different jobs. One is policy analysis for Congress. The other is litigation advocacy in a live case. They are not required to agree, and the gap between them is the honest state of American law right now: unsettled, argued by serious people on both sides, and years from a final answer.

Note

Neither document binds a judge. When someone tells you the law now permits or forbids training on your content, ask which court, which circuit, and whether the decision is final. Most of the time the answer is that no court has said it yet.

A third voice is worth knowing about, because it complicates the tidy two sided story. On 31 August 2026 the Electronic Frontier Foundation filed amicus briefs in two AI copyright cases, arguing against the market dilution theory that rightsholders have been pressing. Its objection is not that AI companies deserve a break. It is that giving copyright owners a veto over works that merely compete with theirs would break the doctrine for everyone, including the small creator that theory claims to protect. You do not have to agree with it. It is a reminder that expanding copyright is not automatically the seller friendly position.

What can a seller actually control?

Less than you would like, and more than nothing. Copyright litigation is not your lever. Your levers are the crawler directives on the property you own and the contracts you sign everywhere else.

Start with the site you control. A robots file can name individual AI crawlers and refuse them, separately from the search crawlers you want to keep. That is a signal rather than a wall: it is respected by the operators who publish a token and honour it, and ignored by anyone who does not. We went through the trade offs of each option in a longer piece on whether to allow, block or charge the AI crawlers hitting your shop, and the short version is that blanket blocking costs you visibility in the AI answers that increasingly sit above the results.

Does blocking AI crawlers hurt your Google ranking?

Not in the case of Google's training crawler. Google documents that Google-Extended is a separate token that governs whether crawled content may be used to train future Gemini models, and states that blocking it neither removes a site from Google Search nor acts as a ranking signal there. Two lines in a robots file, no search cost, is about as clean as a decision gets.

Other crawlers are their own decisions with their own consequences, and they are not symmetrical. Refusing the crawler that feeds an assistant's answers can remove you from those answers, which is a real cost for a shop whose customers now ask a machine where to buy. That trade off deserves a deliberate choice, not a copied file.

Now the part that catches people. If your products live on a marketplace, your robots file is irrelevant to that listing, because you do not own the domain. What governs there is the seller agreement, and those agreements generally grant the platform broad licence to use listing content, including for machine learning. Nothing filed this week changes a contract you already signed. If you sell on a platform and also run your own storefront, the two halves of your catalogue sit under completely different rules, which is one more argument for a storefront whose code and content you actually own. That is the whole design of an AI store builder that leaves the code and the domain in your hands.

Card listing the three practical controls a small seller holds over AI training: crawler directives, marketplace seller terms, and provenance records for assets

Where Europe parts company

If you sell into the European Union, the American argument is interesting rather than decisive, because the legal machinery is built differently. EU copyright has no fair use doctrine. It has a closed list of exceptions, and the relevant one for this purpose is the text and data mining exception introduced by the 2019 copyright directive, which lets rightsholders reserve their rights and opt out.

That opt out is the European equivalent of the lever you were reaching for. It is expressed through machine readable signals on your own site, which is why the robots question above matters more on this side of the Atlantic than it looks. The AI Act then reinforces it: Article 53 requires model providers to respect those reservations, and Recital 106 extends the expectation to training carried out anywhere, so a provider cannot train under a permissive regime abroad and sell the result into Europe untouched. A Munich court has already ruled against OpenAI over memorised song lyrics, a judgment now under appeal.

The practical consequence for a seller with European customers is that reserving rights is a defined act with a defined effect, rather than a hopeful gesture. If you intend to reserve, do it properly and document when you did it. We covered the wider compliance picture in the obligation by obligation map of the AI Act, and the training rules are only one row of it.

The other direction: what you use, not what you make

Everything above treats you as a source of training data. Flip it around, because most small sellers are exposed on the other side too. You use generated images in ads. You use generated copy on product pages. You may have a generated logo.

The output question is untouched by this week's filing, and it was always the riskier one for a buyer of AI tools. A model that reproduces protected expression creates a problem for whoever published the result, and that is you. The defence is unglamorous: keep the prompt, keep the date, keep the tool version, keep the licence terms that applied when you generated the asset. Vendors change those terms and models get retired, so a record made at the time is worth more than a memory of what the page said. If any of the material is going onto packaging or a trademark application, the ownership question gets sharper, and we worked through it in the piece on the AI logo you cannot copyright but can still trademark.

There is a second reason to keep those records. The market is drifting toward provenance signals attached to files, and toward platforms asking sellers to declare when an image was generated. A shop that can answer the question in five minutes will find that drift boring. A shop that cannot will find it expensive.

How would you know if your content was used?

Usually you would not, and that asymmetry is the quiet unfairness at the centre of the whole dispute. There is no register of what went into a training run, no notice sent to the sites that were read, and no practical way for a shop with 400 products to audit a dataset it cannot see.

The nearest thing to evidence available to a small business is behavioural. Ask an assistant a question only your site answers, phrased the way your site phrases it, and see whether the reply carries your specific wording or your specific numbers. That tells you something about what the model absorbed, though it proves far less than it appears to: the model may be reading your live page through a search tool rather than reciting anything it memorised. The two look identical from the outside and are legally quite different, which is precisely why the output question in the table above stays separate from the training question.

If you do find your exact sentences reproduced at length, that is worth a screenshot and a date. It is the kind of concrete instance lawyers actually work with, unlike the general grievance that your work was somewhere in a pile.

What a reasonable seller does about it

Very little, deliberately. The AI training copyright question will not be settled by one filing, and nothing you do this week changes how it ends. This is not a week that demands action, and treating every legal headline as an emergency is how small businesses waste the attention they need for the parts of the year that actually move revenue.

Four things are worth an hour, once, and they hold regardless of how the case ends. Decide which AI crawlers your own domain welcomes and write the file, including the Google-Extended line if you want to keep your content out of Gemini training without touching search. Read the machine learning clause in the seller agreement of whichever marketplace carries most of your revenue, so you know what you granted. Start a plain text log of the AI assets you publish, with dates and tools. If you sell into Europe, reserve rights explicitly rather than assuming silence protects you.

Then stop reading about it. The consolidated New York cases will run for years, appeals included. The Copyright Office may publish more. Courts in Munich and elsewhere will keep producing judgments that point in different directions. None of that has to be tracked by someone whose real problem this quarter is stock, or margin, or the fact that nobody answers the phone on Saturdays.

The department argued that copying works to train a model, considered on its own, should generally be treated as fair use.Justice Department statement of interest, filed 2 September 2026

One last framing that we think holds up better than the headlines. This case is not really about whether your product descriptions are special enough to protect. It is about who has to ask permission before building something at scale, and the answer arrived at will shape what the tools you rent next year are allowed to have learned from. That is worth understanding once. It is not worth watching daily.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building