- Deleting a customer from your shop does not delete them from the AI tools that processed their messages, and you have one calendar month to answer an erasure request.
- Default retention differs by vendor: OpenAI keeps abuse monitoring logs up to 30 days, Google's Gemini API keeps logs up to 55 days by default with 7, 14 and 28 day options, and Anthropic deletes consumer chat data from back end storage within 30 days.
- Zero data retention exists at all three providers, is granted on approval rather than by ticking a box, and never covers everything you use. Setting your own deletion dates matters for the same reason it does in personalisation built on customer order history: data you keep is data you have to answer for.
- The exclusions follow the same rule everywhere: any feature that stores state is outside the arrangement. Files, batches, assistants, threads and vector stores all keep data until something deletes it.
- Flagged content escapes every retention promise. Anthropic states it may hold inputs and outputs for up to two years when trust and safety systems flag a session.
The email is four lines long and perfectly polite. A customer wants their data deleted. You open the admin, find their record, delete it, and reply that it is done.
It is not done. The message they sent your support inbox last spring went through an AI drafting tool. The call they made was transcribed. The complaint they wrote was pasted into a chat assistant so you could work out how to answer it. Each of those handoffs left a copy somewhere, on a clock that somebody else set, and your reply just told a customer something you did not verify.
This is not a reason to stop using the tools. It is a reason to know the numbers, because AI data retention is published, specific and almost never read by the people it binds. Your support ticket archive is where it bites first.
What does an erasure request actually cover?
Less than people fear in one direction and more in another. The right is not absolute, and the Information Commissioner's Office notes that it applies only to data held when the request arrives, not to anything created afterwards. You have one calendar month from the day after receipt to respond.
The part that catches small businesses is the onward obligation. Where you have disclosed the personal data to others, the ICO's guidance says you must contact each recipient and tell them about the erasure, unless doing so is impossible or takes disproportionate effort. Your AI vendors are recipients. So is the model provider sitting behind them.
Nobody expects a two person shop to run an enterprise privacy programme. What is expected is that you know who holds what, which is a list you can write in an afternoon and never have to write again.
Where does one customer message actually go?
Further than the tool you typed it into. A single support message routinely lands in four or five places, and only the first one appears in your admin panel.
The fourth step is the one that surprises people. When you use an AI feature inside a helpdesk or a shop platform, that vendor is usually not running its own model. It is calling one of three or four providers, which makes the provider a subprocessor of your vendor and your vendor a data processor acting for you, and the retention rules that apply to your customer's message are the provider's rules, not your vendor's marketing page.
This is why the question to ask a vendor is not "do you store my data" but "which model provider do you call, and under which arrangement". The answer is in their subprocessor list, which any vendor selling to businesses in Europe is obliged to publish.
How long do the big providers actually keep it?
Between zero and two years, depending on the provider, the feature and whether anything got flagged. The table below is assembled from each provider's own documentation, which is worth reading directly because the numbers move.
| Provider and surface | Default retention | Configurable? | The exception |
|---|---|---|---|
| OpenAI API, most endpoints | Abuse monitoring logs up to 30 days | Zero retention on approval | Conversations, Assistants, Threads and Vector Stores are kept until deleted |
| OpenAI Responses API with storage on | 30 days minimum | Turn storage off | Prompt caching holds data up to 24 hours regardless |
| Google Gemini API, paid tier | Logs up to 55 days | Yes: 7, 14, 28 or 55 days | Logs kept in datasets have no set retention period |
| Anthropic Claude, consumer chat | Removed from back end storage within 30 days of deletion | Not on consumer plans | Flagged sessions up to 2 years, safety scores up to 7 years |
| Anthropic Claude API | Zero retention available by arrangement | Yes, on approval | Files, batches and code execution sit outside it |
Two things jump out once the rows sit together, and neither is visible when you read any single vendor's page on its own.
The first is how close the defaults are. Thirty days at two providers, fifty five at the third, on a control almost nobody adjusts. If you have been assuming that one of the big three is meaningfully more careful with your customers' text than the others by default, the published numbers do not support it. They differ on configurability, not on appetite.
The second is that the exceptions column is doing more work than the default column. Every row has one, and in four of the five cases the exception describes a longer or unbounded retention than the headline figure. A shop that reads only the first number and plans around it has planned around the case that does not apply to the features it is most likely to use.
Why does zero data retention never cover everything?
Because storage is the feature. This is the pattern the table exposes, and it holds across all three providers without exception: the moment a capability needs to remember something between requests, it cannot also promise to remember nothing.
Anthropic states the logic plainly in its documentation, noting that features excluded from zero data retention are fundamentally stateful, because the batch API stores your jobs and the files API stores your files, and that choosing those features is a choice to step outside the arrangement. OpenAI's own eligibility list draws the line in the same place, with assistants, threads, vector stores and conversations excluded while plain chat completions and embeddings are eligible.
For a merchant, that turns a vague worry into a specific question about your own stack. A tool that answers one message at a time and forgets is on the safe side of the line. A tool that builds a searchable memory of your customer conversations, which is exactly what the good ones advertise, is on the other side by construction. The capability you are paying for is the capability that keeps the data.
Zero data retention is not a checkbox at any of the three providers. All of them describe it as subject to approval and additional terms. If a small vendor tells you their AI features run under zero retention, ask whether that arrangement is theirs, and whether it survives the features you actually use.
What happens when something gets flagged?
Every retention promise has a trapdoor, and this is it. Automated safety systems that detect a possible policy violation can hold content well beyond the normal window, and the exception survives arrangements that otherwise guarantee nothing is kept.
Anthropic's documentation states that even where zero retention or HIPAA arrangements are in place, it may retain data flagged by its automated trust and safety systems, with flagged inputs and outputs held for up to two years. Its consumer policy adds that trust and safety classification scores can be kept for up to seven years. OpenAI describes a comparable safety retention path where classifiers detect severe risk.
The practical exposure is small but not zero, and it is strangest in customer service, where the flagging system was designed for a different threat entirely. An angry customer describing a violent thought, a message about self harm, a dispute involving accusations: these are the messages a classifier is most likely to flag, and they are also the messages most likely to be attached to a real named person in your records.
You are not allowed to delete everything anyway
This is the half of the answer that reassures people, and it gets left out of most writing on the subject. A customer asking you to erase everything cannot compel you to destroy records you are legally required to keep.
Invoices and order records sit under tax and accounting law in every jurisdiction a shop operates in, with retention periods measured in years rather than days. The ICO is explicit that the right to erasure is not absolute. Where you have a legal obligation to keep something, the obligation wins, and the correct response is to say so rather than to quietly comply with part of the request.
What that leaves is a genuinely useful split. Transactional records you must keep, marketing consent and profile data you should delete, and the AI layer where a copy exists on somebody else's clock. Those three get three different sentences in your reply, and the third is the only one most shops cannot currently write.
Worth checking while you are in there: what your own systems mean by delete. Plenty of shop platforms soft delete, which hides a record from the interface and leaves the row in place. Backups are the same question with a longer clock, and a backup taken yesterday still contains the customer who asked today. None of that makes you non compliant, because a reasonable retention schedule for backups is accepted practice, but it does mean that the honest answer to a deletion request is a date rather than the word yes.
What does a good reply look like?
Specific, dated and short. The reply that gets a shop into trouble is the one that overpromises, and writing the template once means never improvising it under time pressure.
A workable structure names each category and its outcome. Account and marketing data removed on a stated date. Order and invoice records retained under accounting law, with the period named, and no marketing use. Content processed by service providers deleted according to their published retention, with the longest window named in days. That third sentence is the one this whole article exists to let you write, and once you have the numbers it takes a minute.
Notice what it does not require. It does not require you to name your vendors, to explain your architecture, or to promise a level of erasure nobody in the industry actually performs. It requires you to know your own longest number, which for most small shops turns out to be somewhere between 30 and 55 days, and to say it out loud.
Which of your tools is the actual problem?
Rank them by whether they see identifiable customer text and whether they store it. That ordering puts most of the anxiety in the right place, and it usually clears three or four tools entirely.
Tools that never see a customer's name or message are not part of this problem at all. A model writing product descriptions from your catalogue, an image generator, a forecasting tool reading order quantities: none of these need to appear on your list, whatever their retention policy says.
Tools that see customer text and forget it are a bounded risk with a published clock. A drafting assistant that reads one ticket and returns one reply falls here, and the number in the table above is the whole answer.
Tools that see customer text and build something from it are where the work is. Transcription with a searchable archive, a support assistant with memory across conversations, anything that embeds your tickets so it can retrieve similar ones later. These hold copies indefinitely by design, and deleting the original ticket does not touch the derived copy. That distinction between a record and a thing built from a record is the same one we ran into when sorting information by sensitivity in a one page AI policy built around data classes.
What to do this week
Four steps, none of which needs a lawyer, and all of which pay off the first time a request arrives rather than during it.
- Write the list. Every tool that sees customer names, emails, messages or call audio. Most shops find between four and nine, and are surprised by two of them.
- Ask each vendor two questions. Which model provider do you use, and what is your retention period for content I send. A vendor that cannot answer both quickly has told you something useful. The same instinct applies to the security questionnaire, which we covered in the vendor form your AI supplier passes without trying.
- Turn the dials that exist. Google's Gemini API lets a project cut logging to 7 days instead of 55. OpenAI lets you send requests without storing them. Neither is on by default, and neither costs anything.
- Write down what you cannot delete. If a copy sits in a provider's abuse log for 30 days, that is a fact you can state honestly to a customer. Telling someone their data will be gone within 30 days is accurate. Telling them it is already gone is not.
That last point is the one worth internalising. Answering an erasure request truthfully is easy once you know the numbers, and impossible while you are guessing. The customer asking is rarely trying to catch you out. They want to know it was taken seriously, and a specific answer with a real date reads as exactly that.
Where this is heading
Retention is becoming a competitive feature rather than a legal footnote, which is good news for small buyers. Providers now publish per feature eligibility tables, offer configurable windows, and compete on how little they keep. That was not true two years ago, and it means the answer you got from a vendor last year may be out of date in your favour.
It also means the burden shifts quietly onto you, in the same way that the default AI features your suppliers switched on become your decision the moment you know they exist. When the controls exist and are documented and are free, leaving a log window at its maximum default becomes a choice rather than an accident. Nobody is going to fine a small shop for leaving a default in place. A customer who asks the right question will still notice, and so will the larger buyer running a supplier review on you.
The underlying discipline is the same one that governs every other integration you have connected: know what reaches out of your business, and on what terms. We keep coming back to that in the audit of which AI tools already hold keys to your accounts, and it is why the way a platform handles merchant and customer data is worth reading before you build on it, which is what our security page sets out for anything built with MaShop.