BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/AI customer service: start with the tickets you re…
ToolsAugust 14, 2026
Read · 5 min
ai customer service · customer support

AI customer service: start with the tickets you repeat

Deflection is not resolution. A test for which support tickets a small shop can safely automate, and what the EU now requires you to disclose.

Key takeaways
  • Deflection and resolution are not the same measurement. A deflected ticket only means the customer stopped talking to you, which includes the ones who gave up.
  • Intercom bills a Fin outcome at $0.99, and one of the three billable triggers is the customer not asking for anything further. Silence is priced as success.
  • The widely quoted share of support volume taken by order status questions ranges from 35% to 60% in the same article, with no external source given.
  • Automate by whether you hold the data and whether the answer is identical every time, not by how annoying the ticket is.
  • Since 2 August 2026, Article 50 of the EU AI Act requires people to be told they are talking to an AI, from the first interaction.

The advice you get about automating support is almost always about volume. Find your highest volume ticket type, point a bot at it, watch the queue drop. It is good advice for a company with a support department and a bad heuristic for a business where the support department is you, in the evening, on a phone.

The reason it fails at that size is that your highest volume ticket is usually the one you are worst equipped to answer automatically. Somebody asking where their parcel is wants a fact from a courier's system. If that fact is not already flowing into somewhere a machine can read, no amount of language model is going to conjure it. You do not have a support automation problem. You have a data plumbing problem wearing a support costume.

So the question is not which tickets are most common. It is which tickets you can answer correctly without a human, today, with the systems you already run.

What is the difference between deflection and resolution?

Deflection counts conversations a human never touched. Resolution counts problems that actually got solved. The two numbers look similar on a dashboard and describe opposite outcomes when a customer leaves unhappy.

Lorikeet, which sells support automation and therefore has no commercial reason to say this, puts it bluntly: a deflected or contained ticket means the customer stopped talking to you, and that is not the same as the customer getting what they needed. Somebody who rage quits your chat widget and buys elsewhere is deflected. So is somebody who was helped perfectly. The metric cannot tell them apart.

Now read a pricing page with that distinction in mind. Intercom prices its Fin agent at $0.99 per outcome, and defines an outcome as occurring when the customer confirms the issue is resolved, or when they do not ask for more help after Fin responds, or when Fin completes a workflow including handoffs. The first trigger is a genuine signal. The third is reasonable. The second one is silence.

I am not suggesting anybody is being sneaky. The definition is published, which is more than most vendors manage, and charging per outcome rather than per seat is genuinely fairer for a small shop. The point is narrower and it is about you rather than them: the number you are billed on is the number that flatters the tool. If you want to know whether the thing works, you have to measure something the invoice does not.

Note

The cheapest honest check costs nothing. Once a month, pull twenty conversations the bot closed and read them. Not a sample of the flagged ones, twenty at random. You will know within an hour whether closed means solved on your queue.

Where does the support volume actually go?

Order status, on almost every shop that ships physical goods. The size of that share is where the published figures fall apart, and it is worth knowing they do before you plan around them.

Diagram breaking a small shop support inbox into its main ticket categories, order status, returns, sizing, stock, damaged items and everything else

One widely circulated guide states that order status questions make up 40% to 60% of all inbound ecommerce support contacts, then presents a table in the same article giving the range as 35% to 60%, and cites no external source for either. Other vendor pages put the same figure between 10% and 25%. These cannot all be true of the same population.

What is almost certainly true is the ordering rather than the magnitude. Order status is the largest single category for most shops that post parcels, and it is the most avoidable one, because the customer is asking a question your courier already answered. The honest version of the advice is: measure your own share before you believe anybody's percentage, and the measurement takes an afternoon with a spreadsheet and last quarter's inbox.

Which tickets should you automate first?

The ones where you hold the data and the answer does not vary. That is the entire test, and it produces a different ranking from the volume test.

Card listing the three questions to ask before automating any support ticket type, data availability, answer consistency, and the cost of being wrong

Run every category you handle through three questions. Do I already have the data needed to answer this, in a system something can query? Is the correct answer the same every time, or does it depend on judgement? And what does a confidently wrong answer cost me?

Ticket typeDo you hold the data?Same answer every time?Cost of being wrongAutomate?
Order statusOnly if tracking is wired inYesLow, the customer corrects youFirst, once the plumbing exists
Return policy and windowsYes, it is your own policyYesLow to mediumYes, immediately
Sizing and fitYes, if the spec is publishedMostlyMedium, drives returnsYes, with the spec quoted
Stock and restock datesYes for stock, rarely for datesNo, dates moveMedium, a promise you may breakStock yes, dates no
Damaged or wrong itemNo, needs evidenceNoHigh, this is a upset customerTriage only, never resolve
Refund disputesNo, needs judgementNoHigh, money and reputationNo

The pattern in that table is the useful part. Everything in the automate column is a question about a fact you have written down somewhere. Everything in the do not automate column requires deciding something. A language model is excellent at retrieving and phrasing a fact, and it is structurally unsuited to deciding whether to believe a customer who says the box arrived crushed.

Why does return policy beat order status as a starting point?

Because it needs no integration. Your return window, your condition requirements and who pays postage are facts you already own, written in your own words, sitting on your own site. Pointing an assistant at that text is a configuration job, not an engineering one.

Order status is the bigger prize and the bigger project. To answer it automatically you need order records and courier tracking joined together and reachable by whatever is answering. Plenty of small shops have that data in three places that do not talk. The volume argument says start there. The delivery argument says start where you can finish this week, get a real reduction, and buy yourself the time to do the plumbing properly.

There is a second reason to start with policy questions. They are the ones where a wrong answer is cheapest. If the assistant misstates your return window, a customer emails you and you fix it. If it invents a delivery date, you have made a promise on behalf of a courier you do not control. We went through a similar ordering when looking at the tasks an AI sales agent genuinely cannot take on, and the shape of the answer was the same: retrieval is safe, commitment is not.

What does the assistant actually read?

Whatever you point it at, and that corpus is usually worse than the person choosing it believes. Most small shops discover during setup that their return policy exists in three versions that disagree: one on a policy page, one in the checkout terms, and one in the sentence they have typed into emails four hundred times.

This is the unglamorous half of the project and it is where the quality actually comes from. An assistant grounded in a contradictory knowledge base does not produce contradictory answers in a way you can see. It produces one confident answer, drawn from whichever version it found, and you learn which one when a customer quotes it back at you.

So before configuring anything, do a reconciliation pass. List every place a policy statement lives, including the ones inside your own email drafts and your marketplace listings. Pick the version that is true. Delete or update the others. That work has to happen eventually and doing it first means the assistant inherits a clean corpus rather than an argument.

The same reasoning applies to product facts. If your sizing chart is an image, no assistant can read it, and neither can the shopping surfaces we looked at when Adobe scored product pages as the least machine readable page type on a retail site. Converting that chart to a table in HTML solves a support problem and a discovery problem with one edit, which is the sort of overlap worth looking for when your time is the scarce resource.

What does resolved mean for each category?

Write it in one sentence per type, in your own words, before a vendor writes it for you. The exercise takes twenty minutes and it is the thing that makes every later number meaningful.

For order status, resolved might mean the customer received the current carrier scan and the expected delivery window, and did not write again about the same order. For returns, that the customer was told whether their item qualifies, what the deadline is, who pays postage, and how to start the process, with a link. For sizing, that they were given the actual measurements for the specific variant they asked about, not a general chart. For stock, that they were told the current state and, if it is out, offered a notification rather than a guessed restock date.

Read those back and notice that each one contains a condition about what the customer got, not about what the system did. That is the difference between a definition you can audit and a metric that reports itself. When you later read twenty transcripts, these sentences are the rubric, and disagreements between you and the tool become visible instead of averaging away.

What has to be true before you switch anything on?

Three things, and the third one is new this month. You need a written definition of resolved, you need an escalation path a person actually watches, and since 2 August 2026 you need a disclosure.

Take the definition first. Lorikeet's recommendation is to write down what resolved means for each ticket category, one sentence each, before you look at any tool. Then pull between 100 and 300 historical tickets weighted toward the difficult ones and score the assistant's transcripts against your own rubric rather than the vendor's. For a shop with a small inbox, scale that down: fifty tickets and one honest afternoon still beats a benchmark from somebody else's queue.

The escalation path matters more than the accuracy rate. An assistant that resolves 60% of contacts and hands the rest over cleanly is a good outcome. An assistant that resolves 85% and traps the remaining 15% in a loop is worse than nothing, because those are your hardest cases and now they are also your angriest. Make sure a human can be reached without the customer having to guess a magic phrase.

Do you have to tell customers they are talking to an AI?

In the EU, yes, as of 2 August 2026. The European Commission's guidance on Article 50 of the AI Act states that providers of systems which interact directly with people, naming chatbots, AI agents and avatars, must ensure people are informed they are interacting with AI, and that this happens from the start of the first interaction in a clear and distinguishable manner.

There is an exemption where the fact is obvious, judged against an average person who is reasonably well informed and observant. The Commission says that exemption should be read restrictively, which is a polite way of saying do not rely on it. The information also has to be understandable and perceivable, tied to accessibility requirements, so a grey disclaimer in six point type is not a disclosure.

Practically this is one line in your widget, and the businesses grumbling about it are missing that it helps them. A customer who knows they are talking to software asks simpler questions, escalates sooner when it is not working, and is much less annoyed when it fails than one who thought they were talking to a person named Alex. The wider obligation map for a small seller is in our breakdown of which EU AI Act duties actually land on a small business, because most of the Act does not apply to you and knowing which parts do is the whole job.

How do you tell whether it is working?

Track repeat contact rate, not deflection. If the same customer comes back within seventy two hours about the same order, the first conversation did not resolve anything regardless of how it was recorded.

Three numbers are enough for a small shop. Repeat contact rate within seventy two hours tells you whether closed meant solved. Escalation rate tells you whether the scope is right, and a rate near zero is a warning rather than a triumph, because it usually means the handoff is broken. Time to first human response on escalated tickets tells you whether the automation has quietly made your worst cases slower, which is the most common way this goes wrong.

Notice what is not on that list. Not a resolution percentage, because you now know how much definition work sits behind that number. Not customer satisfaction scores on bot conversations, because the people most annoyed by a bot are the least likely to fill in a survey about it.

What this costs, honestly

Per resolution pricing looks small and is not always small. At $0.99 an outcome, a shop handling 400 support conversations a month is looking at roughly $400 if the tool engages with all of them, against a subscription that might have been a fifth of that. Zendesk's automated resolutions were listed at $1.50 to $2.00 each in the same comparison, on top of seat fees.

Whether that is good value depends entirely on what your own time is worth and how much of it support eats. For somebody spending six hours a week on repeated questions, a few hundred dollars a month is straightforwardly worth it. For somebody spending forty minutes a week, it is not, and the correct move is a decent FAQ page and a tracking link in the dispatch email. Not every problem needs a model.

The other cost is the one that does not appear on an invoice. Every automated answer is an answer you no longer read. Support tickets are the highest quality product feedback a small business gets, and routing them all past yourself has a real cost in things you stop noticing. Keep reading a sample. Whatever else you automate, do not automate away your own knowledge of what customers are confused by.

"The only resolution number worth trusting is one produced from your queue, under your definition."Lorikeet, on AI support resolution benchmarks

Is a chatbot on the site even the right shape?

Often not. For order status the better intervention is proactive: send the tracking link at dispatch, send a note when the courier reports an exception, and the question never gets asked. The reduction methods listed alongside those WISMO figures are all outbound rather than conversational, which is a quiet admission that the best support automation is the ticket that never arrives.

A widget helps when somebody has a question your pages did not answer and they want it now. It is the wrong tool for a question you could have answered before they thought to ask. If you are choosing between building a chat experience and wiring your tracking into your dispatch email, wire the tracking. The capability ordering we set out in the ladder of what conversational systems can reliably do holds here: retrieval before conversation, and notification before either.

All of this assumes you can change your own site. Adding a disclosure line, publishing your return policy as readable text, putting a tracking link into a transactional email: on a hosted platform each of those is a plugin search and a wait. When the storefront code is yours, they are an afternoon, which is the practical argument for running a shop you can edit yourself rather than one you file requests against. The regulation arriving this month is a good example, because everybody with a chat widget needed the same one line change on the same date.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building