- Generating forty ad creatives now takes an afternoon. The arithmetic that decides whether a winner is real did not move at all, and it is the binding constraint for a small advertiser.
- At a 2 percent conversion rate, detecting a 20 percent improvement needs roughly 19,600 visitors per variant. At 300 visitors a day that is 65 days for a two arm test.
- Test eight creatives against a control and you are running seven comparisons. The chance that at least one looks like a winner by luck alone rises to about 30 percent.
- Checking results daily and stopping when something goes green turns a 5 percent false positive rate into 26.1 percent, by the calculation Evan Miller published and nobody has refuted.
- Google states plainly that its Ad Strength rating does not affect serving eligibility, Ad Rank, Quality Score or auction outcomes. It is feedback, not a lever.
- For most small shops the honest move is to stop running tests and start running rotations, judging creative on a longer horizon by whether the account improves.
An AI image tool will give you twelve versions of your product on twelve backgrounds before your coffee is cold. A language model will write thirty headlines. This is genuinely useful, and it has created a problem nobody had in 2019, which is that the supply of creative now vastly exceeds the traffic available to judge it.
The question in the title is not rhetorical. There is a number, it depends on your conversion rate rather than your ambition, and for most small advertisers the answer is smaller than the number of variants an AI tool will happily hand them.
How much traffic does a creative test actually need?
More than people expect, and the shortfall is not close. The standard rule of thumb for a fixed sample test at 80 percent power and 5 percent significance is published on Evan Miller's page about how not to run an A/B test, where he gives the sample size as 16 times the variance divided by the square of the effect you want to detect. For a conversion rate that variance is the rate multiplied by one minus the rate.
Applying that formula to a 20 percent relative improvement, which is a large win by any standard, gives the table below. The visitor counts are per variant, so a two arm test needs double them.
| Baseline conversion rate | Visitors per variant | Days at 300 visitors a day | Realistic for whom |
|---|---|---|---|
| 1 percent | 39,600 | 132 | Nobody running a single shop |
| 2 percent | 19,600 | 65 | A season, for one question |
| 3 percent | 12,900 | 43 | Possible, once a quarter |
| 5 percent | 7,600 | 25 | Workable for a strong offer |
| 10 percent | 3,600 | 12 | Email lists and warm retargeting |
Those numbers are ours, computed from the published rule rather than lifted from anywhere, and they are the reason this article exists. A shop doing 300 sessions a day at a 2 percent conversion rate can answer roughly five creative questions a year with statistical confidence. An AI tool can generate five hundred creatives in that time. The mismatch is the whole story.
Evan Miller's sample size calculator makes the same point interactively, and its wording about a grey zone is worth internalising: conversion rates inside that band are simply not distinguishable from your baseline, however long you stare at the dashboard.
What happens when you test eight creatives at once?
You stop running one test and start running seven, and the mathematics of that is unforgiving. Wikipedia's entry on the multiple comparisons problem states the mechanism directly: each test carries its own chance of a false positive, and the probability that at least one appears grows with the number of tests, because the confidence levels multiply rather than add.
Run the arithmetic for a small advertiser. Three challengers against a control, each judged at the usual 5 percent threshold, gives a 14.3 percent chance that at least one looks like a winner when none is. Seven challengers takes that to 30.2 percent. At eight or ten variants, which is exactly what a generation tool encourages, you are more likely than not to promote a creative that is no better than what you had, and you will do it with a green tick beside it.
The fix is not a correction formula, though those exist and the same Wikipedia entry lists them. For a shop the fix is editorial: pick the one question that matters this quarter and test that. Everything else goes into rotation rather than into the experiment.
Why does daily checking make it worse?
Because the significance calculation assumes you decided the sample size before you started. Miller's article is built around that violation, and the number he reports for it is the one to remember. Monitoring continuously and stopping as soon as a result crosses the line turns a nominal 5 percent false positive rate into 26.1 percent. He also gives the compensating thresholds: after a single peek you would need to see 2.9 percent to claim 5 percent, and after ten peeks 1.0 percent.
This is the single most common failure in small account management, and it does not feel like a mistake while it is happening. You check on Tuesday because you are curious, variant B is up, you pause A to stop wasting money, and you have just concluded an experiment on the basis of noise. Nobody sets out to do this. The dashboard invites it hourly.
Write the sample size and the stop date on the same line as the test name before you launch it. It sounds bureaucratic for a business of one. It is the only intervention on this page that costs nothing and works every time.
Where the platforms already do this for you
Both large ad platforms have moved the allocation problem inside their own systems, which changes what your job actually is. Google's responsive search ads take up to 15 headlines and 4 descriptions and assemble combinations at auction time, and its guidance on creating effective responsive search ads asks for as many unique high quality assets as you can supply rather than for a tested winner.
Two things in Google's own documentation deserve more attention than they get. The first is a correction of a widespread belief: the page on Ad Strength for responsive search ads says the rating does not directly influence serving eligibility, and does not affect Ad Rank, Quality Score or auction wins. It is a feedback tool. Advertisers rearranging headlines to chase a green label are optimising a hint rather than an outcome.
The second is the shape of the reported gains. Google reports that moving Ad Strength from Poor to Excellent is associated with about 15 percent more clicks and conversions on average, that adding a second responsive search ad to an ad group gives 6.6 percent more conversions, and that a third adds 3.7 percent. Read those three numbers together and the diminishing return is visible: the first ad matters, the second helps, the third is small. That is a useful prior for how much creative volume is worth building.
Where AI genuinely earns its place here is in supplying variety rather than in picking winners. Fifteen headlines that all say the same thing in different words give the system nothing to work with. Fifteen that address different reasons to buy do. Our comparison of AI copywriting tools and what each is good at covers which tasks produce real variation and which produce synonyms.
The five levers, and the rule about moving one
Creative is not one thing, and lumping it together is why so many tests are uninterpretable. There are five levers, shown in the figure above, and they fail in different ways, which is also why the choice between generated and filmed footage belongs in its own decision rather than in the test, as we set out in the guide to which AI product video jobs are safe to ship.
- The hook. The first three seconds or the first line. Cheap to vary, high variance, and the lever AI is best at supplying.
- The offer. What you are actually proposing. The highest impact lever and the one people test last, because changing it costs money rather than time.
- The format. Still image, video, carousel, user footage. Often dominates hook and offer combined, and it interacts with placement.
- The proof. Reviews, before and after, a number. Cheap, and frequently the difference between a click and a purchase.
- The audience. Not creative at all, but it is the variable most often changed accidentally in the same week as a creative change.
The rule that follows is boring and it is the whole discipline: move one lever per test. If the winning creative differs from the control in hook, format and offer, you have learned that some combination works and nothing about which part to reuse. Reusability is the only reason to test in the first place.
What should a shop with modest traffic do instead?
Rotate, then judge on the account rather than on the variant. Concretely, that means running four to six genuinely different creatives at once, letting the platform allocate delivery, refreshing the weakest one every few weeks, and asking whether cost per acquisition for the account improved over a period long enough to mean something. You give up the ability to say which creative won. You keep the ability to spend sensibly, which is what you actually needed.
Three habits make that approach work rather than drift.
- Keep a creative log. Date, what changed, what you expected. Six months of that is worth more than any single test, because patterns emerge across campaigns that no individual test has the power to show.
- Judge on money, not on rate. Click through rate is the metric AI creative moves most easily and the one least connected to your bank balance. A hook that lifts clicks and drops conversion rate has cost you.
- Refresh on fatigue, not on schedule. Watch frequency and the trend in cost per result for a given creative. Replacing something that is still working is a self inflicted wound.
Does AI generated creative perform differently?
On the evidence available to a small advertiser, the honest answer is that the generator matters far less than the input, which is the same conclusion we reached about product copy. The failure mode is specific: models produce fluent, generic output when given a generic brief, and generic ad creative underperforms not because a machine wrote it but because it says nothing a competitor could not say.
Two practical guards. Feed the model your actual customer language, taken from reviews and support emails, rather than a description of your brand. And check where a generated image will run before you approve it, because a compelling creative that trips a policy review costs you a campaign start rather than a conversion. Our piece on what happens when an automated ad review disapproves you covers the appeal path when that goes wrong.
There is also a landing page half to this that ad accounts routinely ignore. A 20 percent lift is far easier to find on a product page that answers the buyer's questions than in a headline, and it applies to every campaign at once. The method in our guide to writing product descriptions from your own attribute data is the cheapest conversion work available to most shops, and unlike a creative test it does not need 19,600 visitors to pay off.
Reading a result you already have
Most people arrive at this subject with a test already running and a number on the screen, so it is worth saying what to do with that. Three questions, in order, and the first two usually settle it.
Ask how many conversions each arm produced, not how many impressions or clicks. Conversions are the currency of the calculation, and an arm with eleven conversions cannot support any conclusion regardless of how confident the interface looks. Ask next whether you decided the endpoint before starting. If the answer is no, the confidence figure on screen is not the confidence you have, by exactly the mechanism Miller describes. Ask finally whether anything else changed during the window: a price, a delivery promise, a competitor's sale, a season. A creative test run through the fortnight either side of a holiday is measuring the holiday.
If all three answers are clean, act on the result. If any is not, the result is a hypothesis rather than a finding, which is a perfectly respectable thing to have. Write it down and design the next test around it rather than pretending the current one closed the question.
Why bigger advertisers can test and you cannot
It is worth being explicit about this, because a lot of advice aimed at small shops is written by people describing a different situation. A retailer with 100,000 sessions a day clears the sample requirement for a 20 percent lift in hours, so they can afford to test small effects, run many concurrent experiments, and treat creative as a continuously optimised system. None of that is available at 300 sessions a day, and the gap is not one of skill or tooling.
What scale really buys is the ability to detect small effects. A 3 percent improvement is worth chasing when it applies to millions of sessions and is invisible at any size a single shop operates. That is why the practical advice inverts: a large advertiser hunts small reliable gains, while a small one should be hunting large unreliable ones, meaning changes big enough that they show up without a statistician.
A useful consequence is that the big swing is the right bet for you. Testing a new offer, a different format entirely, or an entirely different promise is more likely to produce an effect your traffic can actually detect than testing two versions of a headline. The counterintuitive part is that this makes small advertisers bolder rather than more cautious.
What about the platforms doing the testing themselves?
They are, and it is mostly good news for a small account, with one caveat that matters. Automated allocation solves the multiple comparison problem by never asking you to declare a winner, and it does so with a great deal more data than your account alone contains. If your objective and your conversion tracking are set up correctly, letting the system distribute delivery across a handful of genuinely different creatives is a better use of your budget than a manual test you cannot power.
The caveat is that you lose the explanation. The system will find a combination that performs and will not tell you why, so nothing transfers to your email, your product page or your packaging. That is an acceptable trade for the ad account and a poor one for the business, which is the argument for keeping the creative log described earlier. The log is where the learning lives when the platform will not give it to you.
Spending the testing budget somewhere it compounds
If your traffic cannot support the tests you want to run, the money is better spent on things whose value does not depend on statistical power: a faster page, clearer photography, a returns policy stated on the product page, an email flow for people who did not buy. None of those need a sample size. All of them lift every campaign you will ever run.
That is also the argument for keeping your fixed costs honest. Before committing a month of budget to creative testing, it is worth knowing what the rest of the stack costs, which is why our pricing page states the build and credit costs plainly rather than asking you to book a call. A test you cannot afford to finish teaches you nothing, and an account with a shorter runway makes worse decisions.
The short version: decide the one question, compute the sample size before you launch, run two arms rather than eight, do not look until the stop date, and put the creative you did not test into rotation instead of into the experiment.