- 404 Media reported on 14 September 2026 that OpenAI pays hundreds of contractors to read real ChatGPT conversations and score the replies.
- Reviewers do not see usernames, but they can see a memory summary that hints at what the account has been used for and roughly where the person is.
- An automated privacy filter runs first, and the company accepts that it misses uncommon identifiers.
- For a private user this is a surprise. For anyone who sells something it is a customer data question, because the useful detail in the box usually belongs to somebody else.
- This is not one vendor behaving badly. Anthropic publishes a thirty day retention window with flagged content read by approved reviewers.
- The practical fix is not abstinence. It is deciding once, in writing, which categories of text never go into a chat box.
A furniture maker I know runs her whole aftersales desk through a chat window. When a customer writes in angry about a cracked table leg, she pastes the email into a chatbot, adds "draft a reply offering a replacement, apologetic but not grovelling", and sends whatever comes back after a light edit. It saves her perhaps forty minutes a day. The email she pastes contains a name, a delivery address, an order number and a complaint.
On 14 September 2026 404 Media published an investigation into an internal OpenAI programme codenamed Project Lily, in which hundreds of paid contractors read real conversations between users and ChatGPT and grade the model's answers. The report was picked up the same day by The Next Web, which framed it as a European data protection problem rather than a curiosity.
Most of the coverage asked what this means for a private person venting about their marriage. That is the wrong question for a merchant. A merchant is rarely pasting their own secrets. They are pasting a customer's.
What did the reporting actually find?
It found a quality process staffed by people. Contractors, recruited through third party staffing firms, are shown a user's prompt and several candidate replies, and asked to summarise what the user wanted and score the options. The stated goal is behavioural: making the assistant less sycophantic and less oddly human. The mechanism is that a stranger reads your sentence.
Two details in the reporting matter more than the headline. The first is that reviewers see whole exchanges rather than isolated fragments, so context accumulates. The second is that above the prompt sits a summary of the account's stored memories, which can reveal what the account has been used for previously and the rough part of the world it sits in. Usernames are withheld. That is a narrower protection than it sounds, because a person is identifiable by what surrounds their name at least as often as by the name itself.
OpenAI runs prompts through an automated privacy filter before a human sees them, and the company has said openly that the filter makes mistakes and misses unusual identifiers. That admission is the honest part of the story and also the operative part. A filter tuned to catch national insurance numbers and credit cards is not tuned to catch "the customer at the veterinary practice on Marsh Lane who threatened to sue us".
Why does this land differently for a shop?
Because the shop is the data controller and the chatbot is not. When a customer hands you their address to receive a parcel, you hold that address under a specific purpose. Pasting it into a third party tool is a transfer, and a transfer is something you are supposed to have told them about.
The Next Web's angle rests on this. It points to a 2025 Court of Justice ruling, EDPS v SRB, which treats the duty to inform as arising at the point of collection and judged from the position of whoever collected the data, not from the position of whoever later receives it. In plain terms, the question is not whether a contractor in another country could work out who your customer is. The question is whether you told your customer that categories of recipient like this exist.
Article 13 of the GDPR sets out what has to be disclosed when data is collected directly from a person, and "the recipients or categories of recipients of the personal data" is on the list alongside the purpose, the legal basis and the retention period. Very few small shop privacy notices mention AI vendors at all. That gap existed before this reporting. The reporting simply makes it concrete, because "a category of recipient" now has a face and an hourly rate.
None of this makes using a chatbot unlawful. It makes using one silently the problem. A privacy notice that names AI assistance as a processing activity, and a habit of stripping identifiers before pasting, closes most of the distance at no cost.
What you paste, and what it would cost if a stranger read it
The useful exercise is not to rank tools. It is to rank the text. Below is the sorting I would use for a small retail or service business, built by taking the categories that actually appear in a support inbox and asking what a reviewer reading them could do with them.
| What goes in the box | Why it is there | What a reader gains | Safer substitute |
|---|---|---|---|
| Full customer email, headers and signature intact | Fastest way to give the model context | Name, address, employer, phone, the dispute itself | Paste the two sentences describing the fault, nothing else |
| Spreadsheet rows of orders for analysis | Asking for a summary or a pattern | A customer list with purchase history attached | Replace names with row numbers before pasting |
| Supplier quote or contract draft | Asking what a clause means | Your cost base and your supplier relationships | Redact the party names and the figures you care about |
| Screenshot of a payment dispute | Asking how to respond | Partial card data, bank references, the customer's claim | Type the question, describe the evidence, attach nothing |
| A staff grievance or reference request | Asking for wording help | Employment details about a named third party | Write it generically, then insert names offline |
Read down the third column and a pattern appears. In almost every row, the sensitive part is not the part the model needs. The model needs the shape of the problem. The identifying detail is carried along because deleting it takes eight seconds and pasting takes one. That is the whole of the risk, and it is an ergonomics problem before it is a legal one. We went through the same sorting exercise in more depth when we looked at which customer data may safely go into an AI tool.
Is this unique to one company?
No, and treating it as a single vendor's scandal is the mistake that leads people to switch tools and change nothing. Human review of some slice of traffic is a normal part of how these systems are built and policed. What varies is the trigger, the retention window and how loudly it is documented.
Anthropic's published retention practice for covered models states a thirty day window for prompts and outputs, with automatic deletion afterwards unless the content was flagged by automated safety systems or has to be kept for legal reasons. Human review there is described as a controlled path triggered by a flag, performed by a limited set of approved reviewers, with access written to logs the reviewers cannot alter. That is a different design from sampling ordinary traffic for quality rating, and the difference is worth understanding rather than cheering.
The honest summary across vendors is this. Some review is triggered by safety flags. Some review is sampled for quality. Enterprise and business tiers generally carry tighter commitments than consumer tiers. Nobody promises that no human will ever see anything. If you are choosing between plans, the thing to compare is the trigger, not the marketing sentence about privacy. Our longer piece on what AI vendors keep and for how long lays out where those commitments are usually written down.
The version of this that a one person business can actually do
Everything above collapses into a short routine. It takes an afternoon once and then costs nothing.
Start by writing down your data classes. Not a policy document, a single page with three headings: text that may go into any chat box, text that may go only into the paid business tier of a named tool, and text that never leaves your own systems. Most shops find the third list is shorter than they feared and contains payment details, anything medical, anything about a named employee, and anything covered by a confidentiality clause with a supplier. We published a template for exactly this kind of one page AI policy because the exercise is more useful than the artefact.
Then change the paste habit. The substitution that saves the most is dull: describe the situation instead of quoting the document. "A customer received a damaged table leg, bought eleven weeks ago, wants a replacement rather than a refund, has written twice" gives a model everything it needs to draft a good reply and gives a reviewer nothing at all. The drafted reply is the same quality. The exposure is zero.
Third, check which tier you are on. Consumer plans and business plans are different products with different commitments, and a large number of small businesses are running company correspondence through a personal account because that is the account they signed up with in 2023. Moving the work to the business tier of the same tool is usually the single highest value change available, and it takes ten minutes.
Fourth, put a sentence in your privacy notice. Something as plain as "we use third party AI services to help draft customer correspondence and to summarise enquiries; these providers act as processors and may retain content for a limited period" does the job Article 13 is asking for. It is not a confession. It is the disclosure that turns an undisclosed transfer into a disclosed one.
What this story does not mean
It does not mean a contractor is reading your messages specifically. Sampling is sampling, and the volume of traffic through a consumer assistant is enormous relative to the number of people employed to rate it. The realistic risk to any individual conversation is small.
It also does not mean the contractors are the threat. They work under confidentiality obligations through staffing firms, and the reporting does not allege misuse by reviewers. The problem the story exposes is structural rather than personal: a data flow that customers were never told about, protected by a filter that its own operator says is imperfect.
And it does not mean you should move everything in house. Running your own model to avoid this would cost more than the risk, and most small businesses would end up with worse security, not better. The proportionate response is smaller than that and always was.
The difference between a shop that has a problem here and one that does not is rarely the tool they chose. It is whether the sensitive half of the sentence was ever typed.
How this connects to the way you build
There is a second order point worth making, because it affects anyone whose storefront or workspace is itself assembled with AI help. The same question applies one level down. When you build a site with an AI tool, where does the tool's context go, what is retained, and who can read it? That is a reasonable thing to ask of any vendor including us, and we document what happens to the contents of a project workspace on our security page rather than leaving it to a marketing line.
The broader habit is to treat every AI surface in the business as a place where text leaves. Your helpdesk macro generator is one. Your marketing copy tool is one. The assistant embedded in your accounting software is one, and it is the one people forget, because nobody thinks of a checkbox in a settings panel as a data transfer.
What if a customer asks whether you used AI on their message?
Answer plainly and have the answer ready before you need it. The version that works is short: yes, we use an AI assistant to help draft replies, a member of staff reads and edits everything before it is sent, and here is what we do and do not send to it.
The reason to rehearse this is that the question now arrives with an edge to it. A customer who has read a headline about strangers reading chat logs is not asking a technical question. They are asking whether their complaint was handled by a person. If your answer is that a human wrote the final message and decided the outcome, you are on solid ground. If the honest answer is that a bot decided their refund and nobody looked, the AI part is the smaller of your two problems.
There is a related trap in the other direction. Do not overclaim privacy. Telling a customer "your message never leaves our systems" when it has been pasted into a consumer chatbot is a false statement about data handling, and it is exactly the kind of statement a regulator treats harshly precisely because it was easy to avoid. Vague is safer than wrong. Accurate is safer than both.
The AI features you did not turn on
The Project Lily reporting concerns a tool people deliberately open and type into. The larger surface for most small businesses is the set of assistants that arrived inside software they were already paying for. A summarise button in a shared inbox. A suggested reply in a marketplace seller app. A writing helper in a spreadsheet. A call transcription feature in a phone system that was switched on by default at the last update.
Each of those is a place where customer text reaches a model, and each carries its own retention and review terms which nobody read, because they arrived as a product update rather than a purchase decision. If you are going to audit anything after this story, audit those rather than the tool you consciously chose. The conscious choice is usually the one with the better terms attached, because you compared something before you bought it.
A workable sweep takes an hour. List the software that touches customer messages. For each one, find the settings panel and look for anything described as AI, smart, suggested or automatic. Note what it reads, not what it produces. Then decide, per feature, whether the convenience is worth the transfer, and write the decision down next to the feature name so that the next person to ask gets an answer rather than a shrug.
What to watch next
Two things. The first is whether European regulators treat disclosure of human review as a distinct failure rather than folding it into the broader training data arguments already under way. The Next Web's reference to the EDPS v SRB reasoning suggests the legal route exists; whether a supervisory authority takes it is a different matter, and Italy's earlier fifteen million euro penalty against OpenAI shows that the appetite is not theoretical.
The second is whether vendors respond by moving the disclosure rather than the practice. The cheapest fix for a provider is a clearer sentence in the privacy policy and a more prominent toggle. That would be a genuine improvement for users who read policies and no improvement at all for the furniture maker, who will still paste the whole email because the whole email is one keystroke and the edited version is four.
Which is why the habit is the part worth changing. Policies move slowly and they move for everyone. The eight seconds it takes to delete a customer's address before you hit paste move immediately, and they move only for you.