MaShop/Journal/Tools/Your Discount Week Sold More. But Did the Offer Do…
● ToolsSeptember 24, 2026
Read · 5 min
incremental sales · holdout test

Your Discount Week Sold More. But Did the Offer Do It?

Revenue during a promotion is not revenue caused by one. How to run a holdout on your own list, and what the biggest published test of this found.

Key takeaways
  • Revenue during a promotion is not revenue caused by a promotion, and no dashboard can tell the two apart because it has nothing to compare against.
  • The largest published test of this, run at eBay across millions of users, found brand keyword ads had no measurable short term benefit and that average returns on non brand search were negative.
  • The cheapest fix available to a small shop is a holdout: keep 10 to 20 percent of the eligible list out of the offer and compare.
  • Check the four weeks after a promotion, not just the week of it. A sale that was going to happen anyway simply moved forward.
  • When you cannot run an experiment, a modelled counterfactual is the honest second best, and it fails in a specific and knowable way.

You ran a discount week. Revenue came in 34 percent above the four weeks before it. The dashboard put a green arrow on the number and attributed the whole thing to the campaign, because that is the only thing it knows how to do.

Here is the uncomfortable part. Some of those buyers were going to buy anyway. They had the tab open. They had been waiting for payday. They got your email, felt pleased about the timing, and used the code on a purchase that was already decided. You paid them 15 percent for the privilege. The dashboard counted every one of those as a win, and it will keep counting them that way for as long as you keep running the promotion.

Separating the sales you caused from the sales you merely witnessed is the single highest value measurement a small shop can learn, and it does not require a data team.

What is incremental sales actually asking?

It asks what would have happened if you had done nothing. That question has no observable answer, because the week where you did nothing did not occur, so every method of answering it is a way of constructing a plausible stand in.

This is the whole difference between prediction and causation, and most business analytics lives on the prediction side without saying so. A model that forecasts next month's revenue from last month's is doing something genuinely useful and is making no claim at all about what happens if you change the price. Those are different questions and they need different evidence.

Comparison figure contrasting what a sales dashboard reports during a discount week against what a holdout test actually measures

How wrong can the naive number be?

Badly wrong, and the best evidence comes from a company with every incentive to find the opposite. eBay ran a series of large scale field experiments to measure whether its paid search advertising caused sales, by switching ads off in some places and leaving them on in others.

The published result is blunt. The paper by Blake, Nosko and Tadelis states that "returns from paid search are a fraction of conventional non-experimental estimates", that "brand-keyword ads have no measurable short-term benefits", and that for non brand keywords, frequent buyers "whose purchasing behavior is not influenced by ads account for most of the advertising expenses, resulting in average returns that are negative".

Read that last clause again with your own shop in mind. The money went mostly to people who were coming anyway. The reason the conventional numbers looked good is that clicks and purchases are correlated: people who are about to buy from you are exactly the people who search for you, click your ad and then buy. Attribution software sees the click before the sale and draws an arrow. The arrow points the wrong way.

A discount code behaves the same way. Your most loyal customers open your emails fastest and redeem first, which means your best customers absorb most of the discount budget while being the group least likely to have needed it.

What does a holdout actually involve?

Withholding the offer from a random slice of the people who would have received it, then comparing the two groups over the same window. Random is the word doing the work, because the moment you choose the excluded group by any rule, you have reintroduced the problem you were trying to solve.

For an email promotion this is genuinely a ten minute job. Most email platforms will split a list randomly. Send to 85 percent, hold 15 percent back, and at the end of the window compare revenue per recipient rather than total revenue, since the groups are different sizes.

The number you want is the difference. If the treated group generated $4.10 per recipient and the holdout generated $3.20, your incremental revenue per recipient is $0.90, and that figure, not the $4.10, is what the promotion earned. Then subtract what the discount cost you on the sales that would have happened regardless, which is the part most people forget.

What you measureWhat it tells youWhat it hidesEffort
Revenue during the promotionThat the week was busyWhether the promotion caused any of itNone
Revenue versus the previous four weeksA rough seasonal comparisonPayday, weather, a competitor going quiet, your own other activityLow
Randomised holdout on the listIncremental revenue per recipient, caused by the offerEffects on people not on the listOne afternoon
Holdout plus four weeks of follow upIncremental revenue net of purchases pulled forwardLong run effects on price expectationsOne afternoon and patience

Why does the week after the promotion matter so much?

Because a discount can move a purchase rather than create one, and the move only becomes visible later. A customer who would have bought on the fourth of next month bought on the twenty eighth of this one. Your promotion week looks excellent. The following fortnight looks mysteriously flat, and by then nobody connects the two.

This is why the follow up window belongs in the measurement rather than in a separate report. Add the four weeks after the promotion to both the treated group and the holdout, and compare the whole period. If the gap that looked like $0.90 per recipient shrinks to $0.30 once the following month is included, you have learned something worth far more than the campaign earned: that this offer mostly rearranges demand.

Note

Run the holdout on the same promotion twice before you act on the result. A single test on a small list carries enormous noise, and the most expensive mistake in this area is killing a campaign that worked on the evidence of one quiet fortnight.

What if you cannot run a holdout at all?

Then you model the counterfactual, accept that it is weaker evidence, and stay honest about the specific way it breaks. This is the situation for anything that hits everyone at once: a site wide sale, a price change, a new homepage.

The standard tool here is a structural time series model that predicts what the treated series would have done without the intervention, using other series that were not treated as its guide. Brodersen and colleagues set this out in the Annals of Applied Statistics in 2015, and the method is the one behind most of the lift figures you will be quoted by an agency.

The failure mode is precise and worth memorising. The documentation is explicit that control series must not "themselves be affected by the intervention", and that the model assumes the relationship between the controls and the treated series established before the intervention "remains stable throughout the post-period". Break either and the method will report a confident number that is wrong in an unknown direction.

For a shop, that translates into a practical rule. Your control has to be something the promotion could not touch. A product category you deliberately excluded works. Total traffic does not, because the promotion drove traffic. A comparable shop's public data works if you can get it, which mostly you cannot.

Card listing four practical ways a small shop can test whether a promotion caused sales without running a formal experiment

Where does AI help, and where does it quietly hurt?

It helps with the arithmetic and the plumbing. Pulling order data into two cohorts, computing revenue per recipient with a confidence interval, and checking whether a gap is distinguishable from noise are exactly the tasks a model with access to a spreadsheet does well and fast. If you can already ask questions of your order data in plain language, this analysis is a few prompts rather than a project, which is the practical side of measuring the return on a single automated task.

It hurts when it is asked to find the answer rather than to compute it. A model handed a sales table and asked what drove last quarter will find a story, because finding stories in tables is what it does, and none of that constitutes evidence about causation. The same applies to any tool that promises to attribute revenue across channels from observational data alone. It is the eBay problem in a new wrapper.

The useful discipline is to decide the comparison before you look at the numbers. Write down the two groups, the window, and the metric, and only then run the analysis. Anything decided after seeing the data is a hypothesis, not a finding.

How does this change what you discount?

Most shops discover the same thing once they measure properly: broad discounts to their existing list are the weakest spend they have, and narrow offers to people who have not bought in a long time are the strongest. The mechanism is obvious in hindsight. A lapsed buyer genuinely might not have come back. A regular was coming regardless.

That reshapes the calendar. Fewer site wide sales, more targeted reactivation, and a habit of excluding recent purchasers from any offer by default. Clearing stock that simply will not move is the separate case, where deciding when to mark down slow stock and by how much is driven by weeks of cover rather than by campaign measurement. It also tends to raise margin without reducing revenue, which is the rare change that needs no trade off. We went through the adjacent question of when a shop may let software move prices at all in the rules around automated pricing, and the measurement discipline here is what stops that automation optimising a number that was never real.

It also changes how you read a vendor's case study. When a tool tells you it lifted revenue 22 percent, the only question that matters is what it was compared against. If the answer is the same shop the month before, you have been shown a correlation with a confident face on it. If the answer is a randomised holdout, the number means what it says. That is the same interrogation we suggested for any AI supplier in the questions to ask before buying.

How big does your list need to be?

Smaller than people assume for a large effect, and much larger than people assume for a small one. That asymmetry is the practical rule, and it tells you which questions are worth testing at your size and which are not.

The intuition is straightforward without any formula. The noise in revenue per recipient comes mostly from a handful of large orders landing in one group and not the other, so the more skewed your order values, the more people you need before a difference means anything. A shop selling one product at one price gets a readable answer from a few thousand recipients. A shop where a single order can be twenty times the median needs far more, or needs to measure order rate rather than revenue.

Two moves make small lists workable. Measure the conversion rate rather than the revenue, because a count is far less noisy than a sum of skewed amounts, and you can convert back to money afterwards using your average order value. And repeat the same test across several campaigns rather than trying to settle it in one, since four modest tests pointing the same way is stronger evidence than one test with a wide interval.

If the honest answer is that your list is too small to detect the effect you care about, that is itself a useful finding. It means the promotion's value has to be argued from margin arithmetic rather than measured, and you should size the discount accordingly instead of pretending a dashboard settled it.

What about a sale that hits the whole site?

Stagger it. If everyone sees the offer at the same moment there is no comparison group anywhere, but very few promotions genuinely have to launch everywhere at once.

The most practical version for a shop is a time stagger inside a single campaign: release the offer to a random half of the list on Monday and the other half on Thursday. Each half acts as the other's control for three days, and you get two readings instead of none. The trade off is that the second group sees a shorter runway, so keep the window identical for both and measure only the first three days of each.

A geographic stagger works if you ship to several countries and your demand patterns are stable enough to compare them, though it is weaker: countries differ in ways that have nothing to do with your offer, and a public holiday in one of them will wreck the reading. A category stagger is usually cleaner. Put the offer on half your catalogue and leave the other half at full price, then check whether the excluded half lost sales, because that number is your cannibalisation and it is often the biggest term nobody calculates.

Whatever you stagger on, write down before launch which group is the control and how long the comparison window runs. A control chosen afterwards is not a control, and the temptation to pick the comparison that flatters the result is strongest exactly when the result matters.

"Because search clicks and purchase behavior are correlated, we show that returns from paid search are a fraction of conventional non-experimental estimates."Blake, Nosko and Tadelis, NBER Working Paper 20171

What should you do this month?

Pick your next email promotion and hold back 15 percent of the list at random. Do nothing else differently. At the end of the window, compute revenue per recipient for both groups, then do it again four weeks later including the follow up period.

Put the result somewhere you will find it again, with the date, the list size, the window and the split. Six of these notes over a year become the only record you own of what actually moves your revenue, and they will outlast any analytics tool you are currently paying for.

You will get one number. It will either be comfortably larger than zero, in which case you now have a defensible reason to keep running that offer, or it will be close to zero, in which case you have just found money you were giving away. Either outcome is worth more than the campaign.

The prerequisite is being able to get at your own order data without asking anyone, which is a reason to care how your storefront stores it in the first place. Our own position on that is that the data and the code should be yours, which is what the MaShop store builder is built around. A measurement habit is only as good as your access to the rows it runs on, and the shops that can answer this question in an afternoon are the ones that never had to ask a platform for permission.

Discussion 0

0 / 4000Your email address is not displayed with your comment.
No comments are published yet.

Explore — related articles.

Build something. Move your work forward.

Start with a software project or an agent task. Describe the result you need, review the work and keep control of your connected accounts.

Open the workspace →