- A randomised experiment across 28 Gap stores found managers attributed unpredictable shifts mainly to wrong shipment information, late promotion changes and head office visits, not to fluctuating customer demand.
- That finding undercuts the usual sales pitch. A scheduler that forecasts demand more accurately is solving a cause that ranked below the ones nobody automates.
- Stable scheduling in that study raised median sales 7% and labour productivity 5%, with an estimated $2.9 million of additional revenue across the treatment stores.
- In Oregon, covered employers must post schedules 14 calendar days ahead, and a late change costs an hour's pay or half rate for hours removed. An optimiser that reshuffles inside the window creates a payroll liability, not a saving.
- Oregon also bars scheduling inside the 10 hours after a previous shift without the employee's agreement, at time and a half if it happens. That is a hard constraint, not a preference to weight.
- Most small teams do not need a scheduler. They need a fixed pattern, a swap mechanism and a rule about who may change what.
The pitch for an AI scheduling tool is always the same shape. It forecasts demand hour by hour, matches staff to it, and hands you a rota that costs less in wages while covering every busy period. It is a good pitch. The evidence suggests it aims at the wrong target.
What actually makes a rota unstable?
Not customers. The Stable Scheduling Study ran a randomised experiment across 28 Gap stores in Chicago and San Francisco between November 2015 and August 2016, with 19 stores receiving the intervention. When researchers asked store managers why shifts changed at short notice, the answers were not about demand. HR Dive's report on the experiment records managers pointing to inaccurate shipment information, last minute changes to promotions and visits by corporate leaders.
Read that list as a small merchant and it translates directly. Your rota does not wobble because Tuesday was unexpectedly busy. It wobbles because the pallet said Wednesday and arrived Thursday, because you decided on Friday to run the sale from Saturday, and because somebody called in sick. Three of those are information problems inside your own business and one is unavoidable.
An AI scheduler improves the demand forecast, which by this evidence is the smallest term. Fixing the delivery information and deciding promotions a week earlier costs nothing and addresses the larger ones. That is an unglamorous conclusion and it is where the money is.
What did stability actually earn?
More than most scheduling software claims to save. The study's results, summarised in the university write up of the stable scheduling experiment, put the median sales lift at 7% and labour productivity up 5%, with an estimated $2.9 million of additional sales over the study period.
The intervention had five parts: a shift swapping app, more consistent start and end times, greater consistency in each person's pattern, a soft guarantee of at least 20 weekly hours for some part time staff, and extra corporate funded staffing in some locations. Before the study even began, the company had already moved all United States stores to two week advance notice and stopped using cancellable on call shifts.
| Change | What it does for the worker | What it does for the business |
|---|---|---|
| Two week advance notice | Childcare and second jobs become possible | Fewer no shows and less churn |
| Consistent start and end times | A life that can be planned | Faster opening routines, less supervision |
| Same pattern week to week | Predictable income | Staff who know the shift they are on |
| Soft guarantee of hours | A floor under earnings | Retention of trained people |
| Shift swap tool | Flexibility without asking a manager | Fewer gaps filled by the owner |
Notice how little of that is a forecasting problem. Four of the five are policy decisions and the fifth is a piece of software that does not need a model at all. A merchant reading a vendor page about optimisation should ask which row of that table the product improves.
Does an automated scheduler create legal exposure?
In some places, directly, and the mechanism catches people out because it looks like efficiency. Oregon runs a statewide predictive scheduling law and its official summary of the requirements sets out obligations that a reoptimising system will violate by design unless it is told not to.
The law covers retail, hospitality and food service employers with 500 or more employees worldwide, so a genuinely small shop is outside it. The reason it still matters is that these rules keep spreading, they apply to franchisees of large brands, and the same shape appears in a growing list of city ordinances. If you are planning a system now, plan it against the stricter rule.
The obligations are specific. Schedules must be in writing at least 14 calendar days before the first day. New hires get a written good faith estimate of median monthly hours and whether on call shifts apply. And any employer requested change inside the notice window carries a cost: one additional hour at the regular rate where hours are added or a shift is moved without losing hours, and half the regular rate for each scheduled hour not worked where hours are cut or a shift is cancelled.
Then there is the rest rule, which is the one an optimiser breaks most naturally. Oregon bars scheduling an employee during the first 10 hours after their previous shift ends unless the employee asks for it, and pays time and a half when it happens. A model minimising labour cost across a week will happily produce a close then open pairing, because on its objective function that is efficient. It is efficient the way skipping insurance is cheap.
These constraints must be hard, not weighted. The difference matters and it is the single most important thing to check in any scheduling tool. A weighted penalty means the system will break the rule when the payoff is large enough. A hard constraint means the schedule is infeasible and it will not be produced. Rest periods, notice windows and contracted hours belong in the second category.
Is stable scheduling worse for margin?
The experiment says no, and that is its most useful result for a sceptical owner. The intuition behind flexible scheduling is that matching labour tightly to demand saves wages. What the Gap stores found was that the tighter match cost more than it saved, through turnover, through slower service by staff who did not know the store, and through the manager time spent filling gaps.
A 5% productivity gain in a small shop looks like an hour of someone's day. A 7% sales lift looks like a customer served properly rather than waiting. Neither shows up in a labour cost report, which is exactly why the optimisation framing wins arguments it should lose. The cost of instability is real and it is recorded in different accounts than the wage bill.
There is a caveat worth stating. That study ran in company owned stores of a large retailer, over ten months, in two American cities, a decade ago. Its mechanism is plausible everywhere and its exact numbers are not a forecast for your shop. Take the direction and the causal ordering, not the percentages.
The three information fixes that come before software
If late rota changes come mostly from bad information rather than from demand, then the highest return work is on the information. Each of these costs nothing and each removes a recurring category of last minute change.
Get a delivery date you can actually plan against. Most small merchants schedule staff against a supplier's promised date and then rework the week when the pallet slips. Ask your two or three main suppliers what their real dispatch to arrival spread looks like, and schedule against the pessimistic end of it rather than the promise. A delivery that arrives early is a pleasant surprise. A delivery that arrives late is two people standing around on Wednesday and nobody available on Thursday.
Decide promotions a week before the rota is published, not after. This sounds trivial and it is the change managers in the study were effectively describing. A sale decided on Friday for a Saturday start guarantees a staffing scramble, and the scramble is invisible in any report because it shows as normal overtime. Put the promotion calendar and the rota calendar in the same place so that publishing one forces you to look at the other.
Put every known interruption in the rota calendar. Stock takes, supplier visits, a landlord's contractor, the accountant's deadline. These are all knowable weeks ahead and each one quietly consumes a person's day. Teams that log them stop being surprised by a category of event that was never uncertain in the first place, only unrecorded.
How do you tell a real scheduling tool from a calendar?
Ask it to break a rule and see what happens. This is a five minute test and it separates products more reliably than any feature list.
Build a week where the only cost efficient answer requires somebody to close at eleven and open at seven. A tool that models rest as a hard constraint will refuse and tell you the week is infeasible. A tool that models it as a preference will produce the rota and perhaps mention it. A tool that does not model it at all will produce the rota silently, and that is the one that will eventually cost you a payment you did not budget for.
Run the same test on contracted hours. Give somebody a 20 hour guarantee and ask for a quiet week. If the system schedules them 12 hours without objecting, it does not understand your obligations and you will be the one who notices, at the point where you are already committed.
The last check is the boring one. Open it on a phone, in a shop, with one hand, wearing the kind of attention you have at four in the afternoon. Scheduling software fails far more often through non adoption than through bad optimisation. A rota nobody looks at is a rota that gets asked about by text message, which is where you started.
What should a five person team actually use?
A pattern, a swap rule and a shared calendar. In that order, before any tool with a model in it.
Fix the pattern first. Write the standing week: who works which days, what time the shop opens, which shifts overlap. Most small teams can express 80% of their rota as a repeating pattern and only argue about the remainder. A pattern also gives you something to measure against, so you can see whether the exceptions are seasonal or just noise.
Decide who may change what, and by when. This is a rule, not software. Staff may swap between themselves up to 48 hours ahead by telling the group. Inside 48 hours it goes through you. Nobody works a closing shift and the next opening. Written down once, this removes most of the negotiation that scheduling tools claim to automate.
Then, if you still need it, buy the tool. What a small team gets from scheduling software is mostly the swap board and the record, not the optimisation. Judge it on whether staff will open it on a phone, and whether it can express your hard rules as hard rules. If it cannot represent a minimum rest gap, it is a calendar with a subscription.
Where does AI genuinely help with staffing?
Two places, and neither is generating the rota. The first is reading your own history to tell you what the pattern should be: which hours are actually busy, how long a delivery really takes to put away, whether the Saturday second person earns their shift. That is analysis of data you already hold, and it is the same argument as the one about restaurant point of sale data in the ranking of which jobs a restaurant should hand over first.
The second is the administrative residue: turning a photographed timesheet into rows, drafting the message that asks who can cover Thursday, summarising a month of hours for payroll. Low stakes, repetitive, checkable at a glance. That is the profile of work these tools are actually good at.
What sits outside both is anything that shades into an employment decision. Ranking staff, predicting who will quit, scoring reliability, deciding who gets the good shifts: those are consequential judgements about people, and they attract obligations that vary by jurisdiction and are tightening everywhere. The reasoning is the same one we set out for hiring tools in the piece on AI resume screening as a compliance purchase. A scoring system aimed at your own staff is a bigger commitment than it looks.
Does better forecasting help at all?
Yes, at the margin, and it is worth knowing when the margin is wide enough to bother. Forecasting earns its keep where volumes are large enough for patterns to be statistically real and where the cost of being wrong is asymmetric. A shop with four staff and eighty customers a day has neither property. A twenty person operation across two sites might.
Choosing the method matters less than most write ups suggest, and we went through how to pick one honestly in the article on choosing a time series forecasting approach. For most merchants, last year's same week plus a manual adjustment for known events beats anything more sophisticated, because the known events are the part the model cannot see.
The order of work
Fix the information that causes late changes: get delivery dates you can trust, decide promotions earlier, put head office visits in the same calendar as the rota. Then publish further ahead than you do now, because the study's own baseline change was two weeks of notice and it happened before the experiment began. Then automate the swap board, which is the piece staff actually want. Only then consider optimisation, and only if your volumes justify it.
Costs are the last question rather than the first. Scheduling tools price per employee per month, which for a five person team lands in the same range as any other small software line, and the discipline for evaluating it is the one in the piece on what a small business actually spends on AI each month: name the job, time it before, time it after. If you want to see how we think about pricing software for businesses this size, our own approach is on the pricing page, and the same test applies to us as to anybody else.
The uncomfortable summary is that the highest return scheduling intervention documented in the research literature required no model at all. It required deciding things earlier and sticking to them. Software can help you stick to a decision. It cannot make the decision earlier, and it cannot fix a delivery date that was wrong before it reached the rota.
Managers said the shifts changed because the shipment information was wrong and the promotion moved. A better demand forecast answers neither.MaShop, on the Stable Scheduling Study, 4 September 2026