BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/AI Inventory Management: What a Forecast Needs Fir…
ToolsJuly 30, 2026
Read · 5 min
ai inventory management · inventory forecasting

AI Inventory Management: What a Forecast Needs First

What a demand forecast needs before it beats a reorder point, what the largest retail forecasting competition measured, and what both stock errors cost.

Twice this quarter you have been out of stock on the thing that sells, while four hundred units of something else sit in the corner of the room where they have been since February. Both errors came from the same source, which is that ordering decisions get made in the ten minutes before a supplier deadline using whatever you can remember.

The obvious question is whether a demand forecast fixes that. The honest answer starts with a number from the largest retail forecasting experiment ever run, and it is not the number the vendors quote.

Key takeaways
  • In the M5 competition on real Walmart sales, only 415 teams, 7.5 percent of those competing, beat the best simple statistical benchmark. The majority were beaten by it by more than 13 percent.
  • Of those 415 teams, just 5 improved on the benchmark by more than 20 percent, and 249 by more than 5 percent. Large gains were rare even among the winners.
  • At the individual product and store level, where a small shop actually operates, the seasonal naive method loses to a method built for sporadic demand by 13 to 27 percent depending on the level.
  • The same paper reports the reverse at aggregate levels, where the seasonal method is on average 11.5 percent more accurate. Which method wins depends entirely on how granular the question is.
  • Intermittent demand, meaning sporadic sales with lots of zero days, is normal at the shelf level and is the condition under which most forecasting fails quietly.
  • The two errors are not symmetric. A stockout costs margin on units you would have sold. Overstock costs cash you cannot spend, and then costs margin again when you discount to clear it.
  • Census figures put AI use in retail trade at 14 percent against a 19.8 percent national average, so nobody in your category has a working forecast either.

What does a forecast need before it beats a guess?

Four things, and most small catalogues fail at least two of them. This is not a criticism of small catalogues, it is a description of what statistical forecasting requires to function.

Enough history, and history of the same thing. A model learns a pattern by seeing it repeat. One year of sales gives you one instance of Christmas, which is an anecdote rather than a seasonality. Two years is the minimum for anything with an annual shape, and three is where it starts to be worth the effort. Once the history exists, choosing what to run on it is a separate exercise, laid out in the guide to the forecasting model ladder and the baseline it has to beat.

Continuity of demand. A product that sells three units on Monday, none for nine days, then seven on a Thursday does not have a demand curve, it has a scatter of events. This is the intermittency problem and it is the norm rather than the exception for individual products in a small shop.

Stability of the thing being forecast. If you changed the price, changed the photograph, ran an ad, or the item went out of stock for a fortnight, that period of history is not a record of demand. It is a record of your own actions. Most small catalogues are mostly this, which is one reason repricing too often quietly ruins the history you forecast from as well as the discounts you can advertise.

A lead time long enough to matter. If your supplier ships in three days, you barely need a forecast, you need a reorder point and an alert. Forecasting pays for itself in proportion to how far ahead you are forced to commit.

What happened when thousands of teams forecast a real retailer's sales?

They mostly lost to simple methods. This is the finding that should reframe the whole topic, and it comes from the M5 forecasting competition, which used 42,840 real series of Walmart unit sales and provided 24 benchmarks to beat, ending in June 2020.

The results are worth reading carefully. According to the competition's results paper, 2,666 teams, 48.4 percent, managed to beat the plain naive benchmark. 1,972 teams, 35.8 percent, beat the seasonal naive benchmark. And only 415 teams, 7.5 percent, beat the top-performing benchmark, an exponential smoothing method applied bottom up. The paper states plainly that the majority of teams were outperformed by that benchmark by more than 13 percent.

Among the 415 who did win, the margins were mostly modest: 5 teams improved on the benchmark by more than 20 percent, 42 by more than 15 percent, 106 by more than 10 percent, and 249 by more than 5 percent. So the distribution of outcomes for people with strong technical skills, competing for prizes, on clean data from the largest retailer in the world, was: nine in ten lose to a standard method, and the winners mostly win by single digits.

Two caveats matter and both cut in useful directions. This was a 28 day horizon on daily data, which is a hard setting. And the data is representative rather than exotic: a companion study comparing the M5 series against data from Corporación Favorita and a Greek supermarket chain found, in its own words on the representativeness of the M5 competition data, only small discrepancies between the datasets.

Note

The lesson is not that forecasting does not work. It is that the benchmark is much stronger than anyone selling forecasting admits, and the benchmark is nearly free. If a simple method applied consistently beat nine out of ten expert attempts, then a simple method applied consistently is your baseline, and any tool has to beat that rather than beat your memory.

Diagram comparing the situations where a simple reorder point beats a demand forecast against where a forecast wins

When does a reorder point beat a forecast?

More often than the software industry suggests, and the conditions are specific enough to check against your own catalogue this afternoon.

Your situationWhat beats whatWhy
Sporadic sales, many zero days per itemReorder point wins clearlyThere is no curve to fit. A forecast on intermittent demand produces a confident fraction of a unit and a false sense of control
Under about two years of historyReorder point winsOne Christmas is not a season. The model will treat a single spike as a pattern and reorder against it next year
New product, or a line you keep changingReorder point, and watch it weeklyNo history exists. Any forecast here is the vendor's category average dressed up as your data
Steady weekly sales on a core line, two years plusA forecast starts to earn its placeThis is the setting the methods were built for, and where even simple exponential smoothing does real work
Strong, repeated seasonality with long lead timesForecast wins decisivelyYou have to commit before you can see demand, so a projection is the only input available. The lead time is what creates the value
Fewer than roughly 30 active SKUsYour own judgement plus reorder pointsAt that size you genuinely can hold the catalogue in your head, and you know things about it no dataset contains

The pattern across the rows is that forecasting value comes from lead time and repetition, not from sophistication. A tool cannot manufacture either.

Why do the granular forecasts fail when the aggregate ones work?

Because the noise that cancels out across a category is exactly the thing you need to know at the shelf, and this is measurable rather than theoretical.

The M5 paper compares Croston's method, built specifically for intermittent demand, against the seasonal naive method across the hierarchy. On average across levels, seasonal naive is 11.5 percent more accurate. But the improvement is 37.8 percent at the highest aggregation level, drops to 9.7 percent at level 9, and goes negative at the three lowest levels, where Croston's is more accurate than seasonal naive by 13.0, 20.2 and 27.0 percent respectively.

Read what that means for you. The lowest levels are individual products in individual stores, which is where every ordering decision in a small shop is made. At that level the method that assumes a repeating seasonal pattern is the wrong tool by a wide margin, and the method that assumes sporadic events is the right one. Any vendor showing you an impressive accuracy figure without saying which level it was measured at is showing you the aggregate number, which is the easy one.

Which kind of AI are we even talking about?

Two completely different things get sold under the same phrase, and confusing them is how merchants end up disappointed by a tool that was never going to do what they hoped.

The first is machine learning applied to demand: gradient boosted trees, neural networks, the family of methods that competed in M5. These are forecasting systems. They need the history and the continuity described above, and their honest performance ceiling is the single-digit-to-teens improvement over a decent benchmark that the competition measured. They do not explain themselves and they do not know anything about your business that is not in the data you gave them.

The second is a language model with access to your stock and order tables. This is not a forecaster and should not be judged as one. What it does is answer questions you would otherwise not bother to ask: which lines have not moved in ninety days, which supplier has slipped on lead time twice this quarter, which products always sell together so a stockout on one kills the other. Those are queries, not predictions, and a language model is genuinely useful at turning a vague question into the right query over your own data.

The second category is where most small shops will get value, and it is not what the phrase implies. Nobody markets it that way because "answers questions about your own spreadsheet" sells worse than "predicts demand", but it works on twelve months of messy history where a forecast does not, and it fails visibly rather than quietly, which is the property you want.

Does holding stock in two places change the answer?

It makes the plumbing matter more than the model, which pushes the decision even further away from forecasting.

With one location, your stock number is either right or you can count it. With two, most of the pain is not knowing what you hold, and the errors compound: a unit shows as available, gets sold, is not there, and the customer finds out after paying. No forecast helps with that. What helps is one number per product per location, updated on every order, with a reorder point per location rather than a single figure for the business, which is a property of the platform holding your catalog, orders and admin rather than of any model bolted onto it.

This is the sense in which stock is an infrastructure problem wearing a forecasting costume. Get to one accurate number and reorder points on top of it, and you will have captured most of the available improvement before any model is involved.

What do the two errors actually cost?

Work it once on one product and you will never argue about this again. Take a product that sells 20 units a month, costs you 12 to buy and sells for 30.

The stockout error. You run out ten days early and lose roughly 7 units of sales. Gross margin on each is 18, so the visible cost is 126. Add whatever share of those buyers went elsewhere and did not come back, which you cannot measure but which is not zero. Call the visible cost 126 and know the real one is higher.

The overstock error. You order 60 instead of 20 and hold 40 extra units. The cash cost is 480 tied up for two months, which matters only if you needed that cash, and small shops usually do. If the line does not move you discount to clear: at 30 percent off you give up 9 per unit across 40 units, which is 360, plus storage and handling and the shelf the units occupied.

So on this product a serious overstock costs roughly three times a serious stockout in cash terms, and the stockout costs more in goodwill. Neither is the small error. What decides which one to lean toward is your cash position rather than a formula: a shop with a cash buffer should bias toward availability, a shop without one should bias toward thin stock and accept the lost sales, because the overstock error can end a business and the stockout error only annoys people.

Card listing the three data conditions a small shop must meet before a demand forecasting tool is worth buying

Does your shop have enough data yet?

Check it in twenty minutes rather than guessing, because the answer is usually specific per product rather than true for the whole catalogue.

Export your sales history by product by week for as long as you have. For each of your top twenty products, count two things: how many weeks of history exist, and how many of those weeks had zero sales. If a product has fewer than 80 weeks of history or more than a third of its weeks at zero, it is not a forecasting candidate today. Set a reorder point for it and move on.

What usually comes out of that exercise is a split: three or four core lines that qualify, and everything else that does not. Which is a genuinely useful outcome, because it tells you to forecast four products carefully and manage the rest with rules, instead of buying a system that claims to forecast all of them and quietly guesses on most.

Retail is not ahead of you on this either. The Census Bureau's measurement of AI use across firm sizes and sectors puts retail trade at 14 percent current use against a 19.8 percent national average, with firms under twenty employees below 20 percent and not moving. Whatever your competitors are doing about stock, it is mostly not this.

What to do this month

  1. Set reorder points on everything. Average weekly sales times supplier lead time in weeks, plus a buffer sized by how bad a stockout is for that line. This is an afternoon and it removes most of the ten-minute-panic ordering.
  2. Pick the four lines that qualify for a forecast. Two years of history, few zero weeks, meaningful lead time. Forecast those with a simple method and compare it against last year's actuals before trusting it.
  3. Measure your own baseline first. Whatever tool you try has to beat your reorder points, not beat your memory. Write down the current stockout and overstock counts before you change anything.
  4. Ask what level any accuracy claim was measured at. Category level or product-store level. The difference is the 13 to 27 percent reversal above, and it is the whole question.

The reason this is unglamorous is that AI inventory management is the one area on a merchant's list where the classical methods are genuinely strong and the data requirements are genuinely strict. It is the opposite of copy, where output is cheap and standards are soft. Our piece on which hours of a founder's week AI actually gives back puts stock decisions in the column where it gives back nothing, and the M5 numbers are why.

Where the software does earn its money is not prediction but plumbing: knowing what you hold, across locations, in real time, with reorder points attached to it and an alert when one trips. That is a data problem rather than a modelling problem, and it starts with getting the numbers off the paperwork at all, which is why pulling line items out of supplier PDFs reliably is usually the first step. It is solved by having stock levels live in the same system as orders instead of in a spreadsheet, which is what an AI generated shop with its own stock and order tables gives you by construction. The credit costs are the number to weigh against one overstock error of the size worked out above, which is a more honest comparison than any accuracy claim. If you are choosing which tool gets a slot at all, the shortlist scored on rewrite rate puts forecasting near the bottom for most shops, and this is the evidence for that ranking.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building