A customer writes in asking for a refund. The message has her full name, her order number, her home address, a sentence about the medication she was taking when the parcel arrived, and the last four digits of the card she paid with. You want to paste the whole thing into a chat window and ask for a polite reply in two languages. Five kinds of data, one paste, and the honest answer is different for each of them.
- Customer data in AI tools is not one question. Sort what you hold into classes first, because the answer changes per class and not per tool.
- The line that decides most cases is contractual, not technical: consumer tiers and business tiers of the same product behave differently by default.
- Anthropic states that by default it does not use inputs or outputs from its commercial products to train models, while feedback you actively submit can be kept for up to five years.
- Google tells Gemini users plainly not to enter confidential information, because a subset of chats goes to human reviewers and reviewed chats can be kept for up to three years.
- Under Article 28 of the GDPR a supplier processing personal data for you needs a written contract, and that obligation exists whether or not the supplier calls itself an AI company.
- Health details, union membership and similar categories inside an ordinary complaint change the answer even when everything around them is routine.
Why is this the wrong question to ask about a tool?
Most people ask whether a given assistant is safe for business use. That question has no answer, because the same brand ships a free consumer product and a business product with opposite defaults, and the difference is written in a contract rather than in the interface. Someone using the free app and someone using the paid business workspace are looking at the same chat box and living under different rules.
The question that does have an answer is narrower. For this piece of information, in this tier of this product, under this contract, what happens to it afterwards. That reframing is the whole job, and it takes an afternoon once.
What are the classes, and which one is your message?
Five classes cover almost everything a small shop touches. The names matter less than the fact that you sort before you paste.
| Class | Example from a shop | Consumer tier | Business tier with a contract |
|---|---|---|---|
| Public copy | Your own product description, your published FAQ | Fine | Fine |
| Internal but not personal | Margin per SKU, supplier lead times, a draft campaign | Commercially unwise | Fine, and this is where most of the value sits |
| Ordinary personal data | Name, email, address, order history | No | Yes, with a signed processing agreement and a privacy notice that says so |
| Special category data | A health detail inside a complaint, religious dietary needs, union membership | No | Only with a specific lawful basis, and usually the answer is to remove it first |
| Credentials and payment details | Card numbers, passwords, API keys, bank details | Never | Never |
That bottom row is not a judgement call and every major vendor says so in its own words. Google's guidance to Gemini users is direct: do not enter login or payment information into the chat. There is no tier, plan or contract that turns a pasted card number into a good idea, because the exposure is not only the vendor. It is your own account, your browser history and whoever else has your screen.
What happens to a message in a consumer tier?
More than most people expect, and the vendors document it. Google's own privacy hub for Gemini Apps states that a subset of chats is reviewed by human reviewers, including trained service providers, with the chats disconnected from the account before they are sent. Default activity retention is 18 months, adjustable to 3 or 36 months. Reviewed chats can be retained for up to three years. Temporary chats, where activity is switched off, stay with the account for 72 hours.
Read those four numbers as a merchant rather than as a privacy specialist. A complaint you pasted in March can still exist in a review pipeline the following spring. If you promised that customer her data would be deleted, and we went through why that promise is harder to keep than it sounds in the piece on what vendors actually keep after you press delete, the pasted copy is a copy you did not count.
The 72 hour figure is the practically useful one. For a genuine one off, where you need help with a single awkward reply and the message contains nothing personal, a temporary chat with activity off is a materially different risk from a chat that lands in an eighteen month history.
What changes in a business tier?
The default on training flips, and that is the headline difference. Anthropic states that by default it will not use inputs or outputs from its commercial products to train its models, naming Claude for Work and its API among them. The same page carries the exception worth knowing: if you explicitly report feedback or a bug, through a thumbs up or thumbs down button for instance, that content may be used for training, and feedback data can sit in secured storage for up to five years.
So the person on your team who clicks the thumbs down button on a bad reply about a real customer has just done something the default protected them from. It is a small mechanic with a five year tail, it appears in no training video, and it belongs in the one page rule sheet described in an AI policy your staff will actually follow.
Defaults change and vendors publish the change on a policy page nobody re reads. Put a calendar reminder twice a year to open the data page of every AI tool you pay for, and diff it against what you told your customers.
Do you need a contract with the vendor?
If personal data goes in, yes, and this is settled law rather than best practice. Article 28 of the GDPR says a controller shall use only processors providing sufficient guarantees to implement appropriate technical and organisational measures. It then requires a contract that sets out the subject matter, the duration, the nature and purpose of the processing and the types of personal data involved, and that binds the processor to act only on your documented instructions, to keep its staff under confidentiality, to respect the conditions for engaging any subprocessor, to help you answer data subject requests, and to delete or return the data when the service ends.
Nothing in that article mentions artificial intelligence, which is the point. A supplier does not escape it by being new. In practice the merchant work is small: find the vendor's data processing agreement, check that it is actually in force for your plan rather than for an enterprise plan you are not on, and keep the PDF where you can find it. If your privacy notice lists the categories of processor you use and this one is not in the list, the notice is now wrong.
Which questions actually separate a good vendor from a bad one?
Four, and none of them is about the model.
- Is my plan covered by your data processing agreement, in writing? Vendors frequently publish one that applies only above a certain tier.
- Do you train on my content by default, and what flips that? Ask about feedback buttons, beta features and anything labelled improve the product.
- Who else sees it, and where do they sit? Human review is normal and rarely advertised. Subprocessor lists are usually public and worth reading once.
- What is the deletion path, and how long does it really take? A delete button in the interface and deletion from backups are different events.
Those four questions do more than a completed security questionnaire, for the reason set out in the piece on the security form your AI vendor passes without trying: a certification proves a process exists, not that your data class is allowed anywhere near it.
What does redaction actually buy you?
Less than the word suggests, and still enough to be worth doing. Replacing a name with a placeholder removes the obvious identifier. It does not remove the order number that maps back to that customer in your own system, and it does not remove the combination of a town, a purchase date and a product that identifies one person in a village.
The European Data Protection Board took a firm line on the general point when it adopted its opinion on AI models and personal data on 18 December 2024. Anonymity has to be assessed case by case, and the test it set is demanding: it should be very unlikely that individuals can be identified directly or indirectly from the model, and very unlikely that their personal data can be extracted from it through queries. The opinion also says that a model built on unlawfully processed personal data can taint the lawfulness of deploying it, unless the model has been duly anonymised.
The merchant lesson is simpler than the legal reasoning. Treat stripping as reducing exposure rather than removing the obligation. If you take out the name and the address and paste the rest, you have made a sensible risk decision, not a legal exemption.
What about the health detail in the complaint?
Take it out before the message goes anywhere. Special category data attracts a higher bar under the GDPR, and the awkward part for a small shop is that it arrives unannounced. Nobody submits a form marked health data. They write that the parcel arrived while they were in hospital, or that they need the fragrance free version because of a skin condition, or that the delivery slot has to avoid a dialysis appointment.
You did not ask for any of that and you still hold it. The practical rule that works in a one person business: when you paste, paste the operational facts and not the story. What was ordered, what went wrong, what the customer wants. The empathy in the reply is yours to write anyway, and it reads better when it is, which is exactly the sort of decision that belongs in the one page AI policy a small shop should write down.
Does any of this apply if you only sell in the United States?
Yes, with different vocabulary and a similar shape. California's framework calls your vendors service providers rather than processors, and the state Attorney General's explainer is blunt about where responsibility sits: it is the business that is responsible for responding to consumer requests, and businesses must tell their service providers to comply with deletion requests. The obligation to reach into a vendor and get data deleted is yours, which means you need to know which vendors hold what before somebody asks.
The same page notes that the right to delete carries real exceptions, including information needed to complete a transaction and information you must keep for legal reasons. That is a useful thing to know before you promise a customer more than the law asks of you, and it is a promise small shops make often, usually in a hurry, usually in writing.
The practical convergence is worth stating. Whether the word is processor or service provider, three things are true in both systems. You stay accountable for data you handed to somebody else. You need to be able to say which vendors hold it. And you need a route to get it deleted that does not depend on goodwill. A shop that can answer those three about every AI tool it uses is in good shape in either jurisdiction, without anyone reading a regulation.
What does a leak look like in a business this size?
Not a breach notification and a law firm. It looks like an old chat history on a laptop that got sold, a shared login used by a seasonal hire who left in January, or a screenshot of a conversation containing a customer's address pasted into a group chat to ask a colleague what to do. The vendor's security is rarely the weak point at this scale. The account boundary is.
Which makes two unglamorous controls worth more than any policy language. Give every person their own login rather than sharing one, so that removing access is a single action when someone leaves. And keep the work account separate from the personal account on the same product, because the personal one usually sits on consumer defaults and nobody notices which window they are typing into at half past nine on a Friday.
Where is the value, if the personal data mostly stays out?
Almost everywhere, which is the part this discussion usually loses. The second row of the table, internal but not personal, is where a small business gets most of the benefit and carries almost none of the risk. Supplier lead times, margin analysis, a pricing structure, category copy, a returns policy rewrite, a plan for next season's range. None of it is personal data. All of it is work you are currently doing at eleven at night.
Support replies are the tempting exception, because the raw material is customer messages. The way through is to work on patterns rather than instances. Take the twenty questions you answer most often, strip them of every identifier, and build your answers from those. That produces something reusable, and it is the approach behind building a help centre out of the support inbox you already have. One pass, no personal data leaving the building, and every future reply gets faster.
Marketing sits in a different place again, because the constraint there is consent rather than confidentiality. What you may do with a customer's behaviour to personalise a shop is a separate question with its own answer, covered in the personalisation you can run without asking permission.
What should a shop actually do this week?
Write one page and put it where your team can see it. It needs four things and nothing else.
A list of your data classes with a real example of each from your own inbox, so nobody has to interpret an abstract category at speed. A named tool and tier for each class, including the classes where the answer is no tool. The two mechanics that quietly change the answer, which are feedback buttons and any new feature that asks to improve the product. And a name, because a rule with no owner stops being true in about four months.
Then check the plumbing once. Confirm your privacy notice mentions the processors you use. Confirm you are on the plan the data processing agreement covers. Confirm that the people who talk to customers know which chat window is which, because the most common failure in a small business is not a policy failure. It is somebody in a hurry, on a phone, using the personal account they signed up with in the first week.
MaShop's own position on all of this is published rather than described, and you can read what we do and do not do with the data that passes through the builder on our security page. Every vendor you use should be able to give you the same thing in the same number of clicks. When one cannot, that is your answer about that vendor, and it did not require reading a single line about how the model works.