MaShop/Journal/Tools/Your Shop Chatbot Just Invented a Product
● ToolsSeptember 29, 2026
Read · 5 min
chatbot · ai assistant

Your Shop Chatbot Just Invented a Product

A bot that invents products was never connected to your catalogue. What a tribunal already decided about liability, and the four records yours has to read.

Key takeaways
  • A tribunal in British Columbia held an airline liable for what its chatbot told a customer, rejecting the argument that the bot was a separate entity.
  • Invented products, prices and policies are not a sign of a bad model. They are a sign the bot was never connected to your actual data.
  • The fix is architectural. A bot that can only answer from your catalogue cannot invent a product that is not in it.
  • Declining to answer is a feature. A bot that says it does not know beats one that guesses, in every commercial scenario.
  • Under the EU AI Act a person must be told they are talking to an AI system, at the latest at the first interaction.

A customer messages your shop at half past nine on a Sunday asking whether the blue one comes in a wide fit. Your assistant answers instantly and cheerfully, and it answers wrongly. There is no blue one. There has not been a blue one since spring.

Every shop that has switched on a chat assistant has some version of this story. The reflex is to blame the model and start shopping for a better one. That reflex is wrong, and acting on it costs money without fixing anything, because the problem is almost never the model's reasoning. It is that nobody told the bot what you sell.

Are you actually liable for what your bot says?

On the evidence so far, yes. The clearest test to date came in February 2024, when the Civil Resolution Tribunal of British Columbia decided Moffatt v Air Canada.

The facts are ordinary enough to be uncomfortable. A customer asked the airline's chatbot about bereavement fares after his grandmother died. The bot told him he could buy full price tickets and claim a refund within 90 days. The real policy required the fare to be applied before purchase. He bought tickets for CA$1,630 and the refund was refused.

Air Canada's defence is the part worth remembering. It argued that the chatbot was a separate legal entity responsible for its own actions. The tribunal member called that submission remarkable and rejected it, finding the airline liable for negligent misrepresentation on the reasoning that the bot was part of its website and the company was responsible for the information there. The award was CA$812 plus costs and interest.

The sum is small and the principle is not. A statement made by your assistant is a statement made by your shop. If yours promises next day delivery to an island you do not ship to, you have made that promise.

Diagram comparing a shop chatbot answering from model memory against one answering from live catalogue data, covering product names, prices and refusals

Why does a bot invent a product that does not exist?

Because it was asked a question about your shop and given nothing from your shop to answer with. Left in that position a language model produces the most plausible answer rather than no answer, and plausible is exactly the failure that gets past a reviewer.

This is the practical, commercial edge of a property we wrote about in why language models hallucinate. A generic assistant knows what shops like yours usually stock, what prices tend to look like, and what returns policies generally say. Asked about a wide fit in blue, it composes something that fits the pattern. Nothing in the process consults your inventory, because nothing in the setup connected it.

The tell is consistency of tone. A bot that is guessing sounds exactly like a bot that is right, which is why shops discover the problem through a customer complaint rather than through testing. Staff who know the catalogue spot it immediately. Customers cannot.

What does connecting a bot to your data actually mean?

It means the answer is assembled from a lookup rather than recalled. The question arrives, the system fetches the relevant records from your catalogue, and the model writes a sentence using only those records.

That pattern has a name in most tooling now and a growing standard behind it. The Model Context Protocol specification describes an open protocol for connecting applications to external data sources, with servers offering three things: resources, meaning context and data; prompts, meaning templated workflows; and tools, meaning functions the model can execute. Your stock levels are resources. Checking a delivery date is a tool.

The specification is also candid about the risks, which is worth reading if you are about to give an assistant access to anything. It states that tools represent arbitrary code execution and must be treated with caution, that tool descriptions should be considered untrusted unless they come from a trusted server, and that the protocol itself cannot enforce its security principles. Connecting a bot to your systems is a permissions decision, not just an accuracy one. We covered the general version of that in what to think about before giving an assistant access to your inbox.

What should a shop bot be allowed to answer?

Questions whose answers exist as data you hold. Everything else should be handed to a person, and the boundary should be drawn deliberately rather than discovered through complaints.

Question typeAnswer lives whereBot may answerFailure if it guesses
Is this in stock in my sizeLive inventoryYes, from the recordA sale you cannot fulfil
What does it cost with deliveryPrice plus shipping rulesYes, from bothA quoted price you must honour
Where is my orderOrder record and carrierYes, after identifying the customerA privacy breach, not just an error
Can I return this after six weeksYour written policyYes, quoting the policyA commitment that binds you
Will this fit my machineCompatibility data, often absentOnly if the data existsA return plus a bad review
Can you do me a discountNowhere, it is a decisionNoA discount your margin did not approve

The last row is the one shops underestimate. An assistant that is agreeable by design will agree to things, and a customer holding a chat transcript offering ten percent has a reasonable expectation. Decisions belong to people, and the bot's job is to route them, which is the same division of labour we described in measuring an AI support agent by resolution rate.

Teaching a bot to say it does not know

This is the single highest value configuration change available to a small shop, and it is usually one sentence in a system prompt plus one behaviour in the retrieval step.

The behaviour is: if the lookup returns nothing relevant, do not answer. Say that you cannot find it and offer the human route. Most people building these assistants optimise for the opposite, because a bot that always answers feels more capable in a demo. In a shop it is a liability generator.

Test it deliberately. Ask your assistant about a product you have never sold, in the tone a real customer would use. Ask about a policy you do not have. Ask for a delivery date to somewhere you do not ship. A bot that produces confident answers to all three is not ready, regardless of how well it handles the questions you designed it for.

Card listing the four live data sources a shop assistant must read, covering stock by variant, current price, delivery estimates and returns policy text

Must you tell people they are talking to a machine?

In the European Union, yes, and the timing is specified. Article 50 of the EU AI Act requires that people are informed they are interacting with an AI system, unless that is obvious to a reasonably well informed person, and the information has to be provided at the latest at the time of the first interaction.

The same article carries obligations around synthetic content, requiring generated audio, image, video or text to be marked in a machine readable format where technically feasible, with carve outs for assistive editing and for content that does not substantially alter the input. For a shop the practical read is narrower than the legal text: label the assistant, and do not let it impersonate a named member of staff.

The commercial argument points the same way as the legal one. Customers who know they are talking to software phrase questions more literally and escalate sooner, which reduces the class of failure where somebody spends twenty minutes negotiating with a program. We went through the labelling detail in what the disclosure rules require of a chatbot.

Note

An assistant that answers from your catalogue has a second use most shops miss. Every question it could not answer is a gap in your product data, written in a customer's own words. That log is the cheapest product research a small shop will ever get.

What does the data side actually require?

Less than a rebuild, more than a spreadsheet. The assistant needs to read four things reliably, and most shops have three of them in a usable state already.

Stock, per variant rather than per product, because the customer is asking about the blue one in a wide fit and an aggregate number cannot answer that. Price, current and including whatever promotion is running, since a bot quoting last week's price has made an offer. Delivery, as rules rather than as a sentence on a page, because the honest answer depends on destination and on what is in the basket. And the returns policy as text the bot quotes rather than paraphrases, since paraphrasing a policy is how a policy changes by accident.

Compatibility and specification data is the common gap. A shop selling parts, consumables or anything fitted to something else is asked compatibility questions constantly, and the answer usually lives in a supplier PDF or in the owner's head. No amount of model quality substitutes for that. Either the data gets captured or the bot hands those questions to a person, and pretending otherwise produces returns.

If you are building the storefront rather than bolting an assistant onto one, this is much easier to arrange from the start, because the assistant and the shop read the same records by construction. That is the approach behind the MaShop MCP server, which exposes a project's real data through the same protocol rather than leaving an assistant to infer it.

How do you tell whether yours is guessing?

Sample the transcripts weekly and check a dozen answers against your own records. Not the flagged ones. A random dozen, because the failures that matter are the ones nobody complained about.

Score each answer on one question: could this have been derived from data we hold? An answer that could not have been, however reasonable it sounds, is an invention that happened to land close. Those are the ones that reveal the architecture problem, and a shop that finds two in twelve has a system that is guessing routinely and getting away with it most of the time.

Watch for a second pattern too, which is drift between the bot's account of a policy and the policy itself. Policies change. Assistants configured with a copy of the text at setup time keep answering from that copy long after the page was edited. Anything the bot quotes should be read at the moment of the question rather than remembered from installation day.

What does this cost a shop that ignores it?

The direct cost is small and the indirect cost is not. A wrongly quoted price on a nine pound item is nine pounds. A wrongly quoted stock level on a wedding order is a review that stays on your listing for years.

There is a pattern in which errors hurt most, and it follows urgency rather than value. A customer buying a gift with a date attached, a part for a machine that is currently broken, a replacement for something lost in transit: these people are the least tolerant of a wrong answer, because a wrong answer costs them time they do not have. They are also, awkwardly, the customers most likely to use chat instead of email, precisely because they are in a hurry.

The second indirect cost is staff time spent unpicking. An assistant that promises something you cannot deliver does not end the conversation, it starts a harder one, and that conversation lands on the same person the bot was meant to free up. Shops measuring their assistant on deflection rate alone miss this entirely, because a deflected question that returns as a complaint was counted as a success.

Track a simple ratio for a month: conversations the assistant closed, against conversations that reached a person after the assistant had already answered. The second number is where the invented answers live, and it is usually higher than anybody expects.

Does a better model fix it?

Marginally, and not in the way the marketing implies. A stronger model hedges more carefully and is less likely to fabricate a confident specification, which lowers the rate without changing the mechanism. It still has no access to your stock.

This matters because upgrading is the expensive response and reconnecting is the cheap one. A shop paying more per conversation for a better model, while still leaving the assistant to guess at inventory, has bought a more articulate guess. The same money spent on exposing the catalogue properly removes the entire class of error.

There is one place where model choice genuinely matters, which is instruction following under pressure. An assistant told to refuse when the lookup returns nothing has to actually refuse, including when a customer pushes back three times. Weaker models fold under insistence and produce the answer the customer clearly wants. That is worth testing directly before you commit, by playing the awkward customer yourself for ten minutes.

The other model dependent factor is language. If you sell across borders, a bot that answers accurately in English and approximately in the other languages you trade in has a silent failure mode, because nobody on your team reads the transcripts in Dutch. Sampling has to cover every language the assistant speaks, and the awkward truth is that most shops only check the one they can read.

A setup checklist for a shop of one

Write down what your assistant is allowed to answer, as a list, before you configure anything. Five categories is usually enough, and the act of writing the list surfaces the questions that have no data answer.

Connect stock and price first, because those are the two that create obligations when they are wrong. Delivery rules next. Policy text last, read live rather than pasted in, so that editing your returns page edits what the bot says.

Then write the refusal. One sentence that names what the assistant could not find and offers the human route with a real response time attached, since offering to pass it to a person is only reassuring if the person appears. Chatbot wrong answers usually survive because the alternative was configured as a dead end.

Finally, put the weekly transcript sample in your calendar for a fixed day. Ten minutes, a dozen conversations, checked against your own records. Everything else in this piece is setup that happens once. That review is the part that keeps working, and it is the part shops abandon after a fortnight.

The short version

Your assistant speaks for your shop and a tribunal has already said so. It invents things when it has nothing of yours to read, which is an architecture problem rather than a model problem, and it is fixed by connecting it to stock, price, delivery and policy, then allowing it to refuse. Tell people it is a machine. Log what it could not answer. Read a sample every week. That is the whole job, and none of it depends on which model you picked.

Discussion 0

0 / 4000Your email address is not displayed with your comment.
No comments are published yet.

Explore — related articles.

Build something. Move your work forward.

Start with a software project or an agent task. Describe the result you need, review the work and keep control of your connected accounts.

Open the workspace →