- The tools that sell AI customer segmentation publish their own minimum data requirements, and most small shops sit under them.
- Google Analytics needs at least 1,000 returning users who triggered the predicted behaviour and 1,000 who did not, inside a seven day window in the last 28 days, before it will produce a purchase or churn probability at all.
- Klaviyo asks for 500 customers who have ordered, 180 days of order history with orders in the last 30 days, and some customers with three or more orders.
- A model that loses eligibility does not announce it. The audience simply stops gathering new people while the dashboard still looks alive.
- Under those floors, rules built from your own order table beat any prediction, because they are right about the past rather than uncertain about the future.
- Grouping customers to decide what to send them is profiling in the legal sense, which is fine, but it carries a right to object that you have to tell people about.
Every email tool now offers to sort your customers for you. The pitch is consistent: stop guessing which buyers matter, let the model find the groups, send each group the right thing. It is an attractive offer for a shop run by one or two people, because segmentation is genuinely tedious work and it is genuinely valuable when done.
What almost nobody mentions is that these features have a floor. Not a recommendation, a requirement, published by the vendors themselves in their own documentation, and a shop below it gets nothing. Not a worse prediction. Nothing. The feature stays greyed out or the audience stays empty, and the owner concludes the tool is broken.
So before any of the strategy, the arithmetic. Find out whether your shop clears the bar, because that single answer decides which half of this article applies to you.
What data do these tools actually require?
Two examples, both documented publicly, both worth checking against your own numbers before you pay anyone.
Google Analytics generates what it calls predictive metrics: purchase probability, churn probability and predicted revenue. The prerequisites page for predictive metrics is unusually direct about the condition. Over a seven day period within the last 28 days, at least 1,000 returning users must have triggered the relevant behaviour and at least 1,000 must not have. Model quality then has to hold up over time for the property to stay eligible.
Read that as a shop owner rather than an analyst. One thousand returning users who bought inside a week is not a small shop. It is a business doing serious volume. A store with 40 orders a month is several orders of magnitude away, and no amount of configuration closes that gap.
Klaviyo sets a different and much lower bar, which is why it is the more relevant example for most independent stores. Its documentation on predictive analytics lists the conditions for the predictions to appear on a customer profile: at least 500 customers who have placed an order, where the orders are not cancelled or refunded and are not zero value; at least 180 days of order history with orders inside the last 30 days; and at least some customers who have placed three or more orders. It then rebuilds the lifetime value model from your data and retrains it at least weekly.
That last condition is the interesting one and it gets skipped. A shop can easily have 500 customers and 180 days of history while having almost nobody who has bought three times. If your product is a mattress or a wedding cake, repeat behaviour barely exists, and a lifetime value model has nothing to learn from.
| Approach | What it produces | Published minimum | Fails by |
|---|---|---|---|
| Google Analytics predictive metrics | Purchase probability, churn probability, predicted revenue | 1,000 users triggering the behaviour and 1,000 not, in a seven day window inside 28 days | Staying ineligible, or losing eligibility silently |
| Klaviyo predictive analytics | Predicted lifetime value, churn risk, expected next order date | 500 customers with valid orders, 180 days of history, orders in the last 30 days, some buyers with three or more orders | Blank fields on profiles with too little individual history |
| RFM rules on your own orders | Named groups you can act on today | Any order history at all | Describing the past confidently and the future not at all |
| First order month cohorts | A retention curve by acquisition month | Any order history at all | Telling you nothing about an individual customer |
Why does a prediction fail quietly rather than loudly?
Because eligibility is evaluated continuously and nothing in the interface is designed to tell you it lapsed. Google's documentation notes that predictive audiences already exported to linked advertising accounts stop accumulating new users if the property becomes ineligible and new predictions are not generated. The audience still exists. It still has a name, a size from before, and a place in your campaign setup.
That is the failure mode worth fearing, and it is specific to the model based approach. A rule never becomes ineligible. If your rule says a customer who bought twice and has not returned in 120 days belongs in a win back group, that rule produces the same answer next January as it does today, whether your traffic tripled or halved.
Seasonality makes this concrete. A shop that sells garden furniture clears a volume threshold in May and falls under it in November. The predictions are strongest exactly when the business needs them least and absent during the quiet months when a win back campaign would earn its keep.
Before subscribing to anything that promises AI customer segmentation, run one query against your order table: how many customers have placed three or more orders, and how many orders landed in the last 30 days. Those two numbers decide whether the feature can work for you. They take a minute to find and they are not on any vendor's landing page.
What should a shop under the floor do instead?
Build the groups by hand, once, and then never think about them again. Recency, frequency and monetary value is the old method for a reason: it needs only the order table you already own, and every group it produces can be explained to whoever writes the emails.
The practical version for a small catalogue is five groups rather than the textbook twenty seven. New buyers who have ordered once in the last 60 days. Repeat buyers with two or more orders and recent activity. Lapsing buyers who used to be regular and have gone quiet past their normal gap. High value buyers in your top decile by spend. Everyone else, which is usually most of the list and should receive the least.
Each group gets exactly one decision attached to it, and this is the part that most segmentation work skips. A group with no decision attached is a report, not a segment. If you cannot say what the lapsing group receives that the repeat group does not, you have sorted your customers for no reason.
The gap between your typical reorder interval and a customer's silence is the only genuinely predictive thing most small shops own, and it costs nothing to calculate. If your median customer reorders every 70 days, someone at 140 days is behaving differently from someone at 75. We went through the probability version of that question in how a small shop can estimate whether a customer is still active, which is the right next step once the grid stops answering.
Does AI add anything a rule cannot?
Yes, in two specific places, and it is worth being precise about them rather than accepting the general claim.
The first is unstructured text. A rule cannot read 400 reviews and tell you that the complaints about sizing cluster around one supplier. A language model can, and that is a segmentation task in everything but name, because it groups customers by a property that never appears in your order table. This is the strongest genuine use for a small merchant and the one least often sold as segmentation.
The second is interaction effects at scale. Once a shop has thousands of buyers and dozens of categories, combinations matter in ways a human cannot enumerate: the people who buy category A and then B behave differently from those who buy B then A. Below that size the combinations are countable and you already know them, because you packed the boxes.
What AI does not add is certainty about an individual. A predicted lifetime value of $99 is a central estimate with a wide spread around it, and treating it as a fact about that person is the most common way these tools get misused. Klaviyo's own field definitions describe predicted lifetime value as a prediction of spend over the next year, sitting beside a historic figure that is simply arithmetic on past orders. The historic number is true. The predicted one is an estimate.
What does a percentile audience actually guarantee?
Its size. Not its quality. This is the single most useful mechanical fact about the predictive audiences a small shop is most likely to switch on, and it is stated plainly in Google's documentation without anyone drawing the conclusion.
The suggested audience called likely seven day purchasers includes users whose purchase probability sits at or above the 90th percentile. Google's own worked example spells out what that means: if the model ran on 1,000 users, the 90th to 100th percentile contains the top 100 of them. The default configuration for churn probability is wider, taking the 80th to 100th percentile, which is the top fifth of users, and you can widen it further to the 40th percentile if you want more people in the group.
Notice what is missing. A percentile is a ranking, not a level. Your top 10% by purchase probability exists whether those users are 80% likely to buy or 3% likely to buy. The audience will always be populated, always be roughly the same share of your traffic, and never tell you by its size whether anyone in it is actually close to buying. Google says as much about widening the range: more users are included, and a greater portion of them are less likely to meet the condition.
For a shop deciding where to spend an advertising budget, that distinction is the whole ballgame. A ranked top decile is a reasonable place to aim a retargeting budget. It is not evidence that the campaign will convert, and the audience filling up is not a signal of anything except that the model ran.
What do the predicted fields on a customer profile mean?
Worth separating, because two numbers that sit next to each other on the same screen have completely different standing. Klaviyo's documentation defines historic lifetime value as the total value of a customer's previous orders with refunds and returns taken into account, and predicted lifetime value as a prediction of what that customer will spend over the next year. Total lifetime value is simply the two added together.
The first of those is arithmetic. The second is a forecast. Adding them produces a figure that is part fact and part estimate, and it is the one most often pasted into a spreadsheet and treated as a customer's worth. If you are ranking customers to decide who gets a hand written note in the box, rank by the historic number. It is the one you can defend.
The same file carries churn risk, an expected date of the next order, and a predicted gender. That last field deserves a separate thought. An inferred attribute about a person is a different category from a record of what they bought, both in how wrong it can be and in how it feels to the customer when it shows up in a subject line. Behaviour you observed is safer ground than a characteristic you guessed.
Is segmenting customers legally sensitive?
It is profiling, and profiling has rules, though probably not the ones you are worried about. The ICO's guidance on rights related to automated decision making including profiling draws the important line: the heavy restrictions apply to decisions made solely by automated means that produce legal or similarly significant effects. Sorting buyers so the lapsing ones get a different email is not that.
What does apply is the right to object to profiling, and the obligation to bring that right specifically to people's attention rather than burying it. In practice that means your privacy notice says you group customers by purchase behaviour to decide what marketing they receive, and there is a working way to say no that does not require unsubscribing from everything.
The line gets closer when segments start driving prices or access rather than messages. A group that receives a different offer is marketing. A group that is quietly excluded from a discount everyone else sees is closer to a decision about a person, and the further you travel in that direction the more the consent question matters. We set the boundaries out in more detail in what personalisation is allowed to use without consent.
How do you know the segmentation worked?
By holding a group back, which almost nobody does and which is the only honest measurement available to a small shop.
Pick your lapsing group. Send the campaign to 90% of it and nothing at all to the remaining 10%. Wait one purchase cycle. If the treated group bought at a higher rate than the untreated one, the segment earned its existence. If the two look the same, those people were going to buy anyway or were never coming back, and the campaign was theatre.
This matters more with predictive tools than with rules, because a model trained on your own history will confidently identify the customers most likely to buy, and sending them an offer they did not need is how a shop discounts revenue it already had. High purchase probability means likely to buy, not likely to be persuaded. Those are different populations and only a holdout separates them.
Run the test once per quarter rather than continuously. A small shop does not have the volume for permanent experimentation, and a holdout that never ends is just a group of customers you have decided to ignore.
Where the data actually comes from
One structural point that decides everything above. Every threshold in this article is measured against clean order data with identifiable customers, and the most common reason a shop cannot segment is not volume but plumbing. Guest checkouts with no account, orders spread across a marketplace and a website with no shared identifier, a platform that hands you a CSV export rather than a queryable table.
If you own the schema, this is a solved problem. The order table has a customer identifier, the queries take minutes, and the thresholds are something you measure rather than guess at. That is one of the arguments for owning the code behind the storefront rather than renting it, which we set out in what an AI built store gives you over a hosted storefront.
Start with the count. Customers with three or more orders, orders in the last 30 days, months of history. Three numbers, one query, and an honest answer about which tools can help you at all.