BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/Customer review analysis: what AI reads, and what …
ToolsAugust 14, 2026
Read · 5 min
customer review analysis · product reviews

Customer review analysis: what AI reads, and what it misses

Reading your reviews with a model beats scoring them. What to extract, what sentiment analysis gets wrong, and the rules on incentives and disclosure.

Key takeaways
  • Reading your reviews with a model is useful and unregulated. Writing reviews with one is named directly in the FTC rule as prohibited conduct.
  • Google will not show star snippets for reviews about a business collected on that business's own site, even through a third party widget.
  • EU traders displaying reviews must publish how they verify them, whether every review gets published, and how the average is calculated.
  • A single sentiment score hides the useful part. Sarcasm, negation scope and mixed opinions all collapse into one misleading number.
  • The most valuable output is not sentiment at all. It is the list of expectations your product page set and the product did not meet.

A shop with four hundred reviews is sitting on the best product research it will ever get, written by people who paid for the privilege of giving it. Most owners read the one star ones, feel bad, and close the tab. The four star reviews, where somebody liked the thing and mentioned one irritation in passing, go unread entirely, and that is where the money is.

Reading four hundred reviews properly is a day's work nobody has. This is one of the few jobs where a language model is straightforwardly good: it does not get bored, it does not skim, and it has no ego about the product. The question is what you ask it for, because the default answer, a sentiment score, is close to useless.

Why is a sentiment score the wrong output?

Because it compresses the only information you needed into a number you already had. You know your average star rating. A model that reads every review and tells you customers are 78% positive has spent tokens reproducing your rating.

It is also the output most likely to be wrong, for reasons that are well documented and structural rather than fixable by a better model. Sentiment analysis has four classic failure modes, and reviews trigger all of them.

Diagram naming five ways sentiment scoring misreads product reviews, including sarcasm, negation scope, ambiguity and mixed opinions in one review

Sarcasm inverts the words. A review saying this phone has an awesome battery back-up of 2 hours uses entirely positive vocabulary to file a complaint. Negation is worse than it sounds because the scope is ambiguous: in "I do not call this film a comedy movie", working out how far the negation reaches is not a solved problem.

Ambiguity is the one that bites product categories specifically. The same source gives the clean example: a story that is unpredictable is being praised, a steering wheel that is unpredictable is a safety complaint. Polarity is a property of the word in your category, not of the word. Nobody can hand you a universal list.

Multipolarity is the one that matters most for a shop, because it is the normal case rather than the exception. "The audio quality of my new laptop is so cool but the display colors are not too good" is two verdicts about two components, and any single score for that sentence throws away the actionable half. The recommended treatment is to pull out each aspect with its own label and only compute an overall figure if you actually need one.

What should you ask a model to extract?

Aspects, expectations and vocabulary. Three passes over the same pile, each producing something you can act on this week.

Card listing three useful outputs from analysing customer reviews, the aspect praised or blamed, the wrong expectation, and the customer vocabulary

Aspects. For each review, what component or attribute is being discussed, and what is the verdict on that specific thing. Sizing, packaging, delivery, durability, colour accuracy, instructions, smell, noise. The output is a table of aspects with counts, and it usually contains one surprise: a component you never thought about generating a third of the complaints.

Broken expectations. This is the highest value question and almost nobody asks it. For every negative or mixed review, what did the customer expect that turned out not to be true, and where would they have got that expectation? Colour looked different in the photograph. Assumed batteries were included. Thought it was waterproof rather than water resistant. Every one of those is a defect in your product page rather than your product, and every one is fixable in an afternoon with no manufacturing change.

Vocabulary. The words customers use when they describe the item to somebody else. These are almost never the words in your product copy, because you write like a seller and they write like a buyer. This list is directly useful for the description, and it is what an assistant summarising your category will match against.

What you ask forWhat you get backWhat you do with itHow often to run it
Overall sentimentA number close to your star averageVery littleNever, you have this already
Aspect and verdict per reviewRanked list of praised and blamed componentsFix the top complaint, promote the top praiseQuarterly, or after a product change
Broken expectationsSpecific false beliefs and where they came fromRewrite the page section that caused each oneQuarterly, and before any big campaign
Customer vocabularyThe words buyers use, ranked by frequencyProduct copy, headings, internal search synonymsTwice a year
Return reason clusteringWhich complaint predicts a returnDecide what to fix versus what to discloseQuarterly

The last row deserves its own note. Complaints and returns are not the same population and you should not treat them as one. Some irritations annoy people who keep the product. Others send it back. If you can join your review text to your return records, even roughly, you learn which complaints cost you money and which merely cost you a star. Those get different responses and different budgets.

Does the star rating tell you anything the text does not?

Sometimes, and the interesting case is when the two disagree. A four star review whose text is a catalogue of problems, or a two star review whose text says the product is excellent but delivery was late, are both worth reading closely.

The second pattern is common enough to distort a whole product's rating. Customers frequently punish the item for a courier's failure, because the star box is the only lever in front of them. If your aspect extraction shows that a large share of your low ratings talk about delivery rather than the product, you have a logistics problem being recorded as a quality problem, and the fix is nothing to do with the item. The same proactive shipping communication that reduces support volume, which we went through in the piece on which support tickets to automate first, tends to move this number too.

How do you ask for reviews without breaking the rules?

Ask everyone, ask the same way, and never mention the rating you would like. Those three constraints keep you clear of both regimes and, awkwardly for the shops that game it, produce a more useful dataset as well.

Asking everyone matters more than it sounds. The common tactic is to send a satisfaction question first and only invite a review from the people who answer well. That filters your reviews into a shape that flatters you, and it is exactly the pattern the EU rules address when they require you to say whether all reviews are published and how the average is computed. It also destroys the analysis this article is about, because the customers who were disappointed are the ones carrying the information you need.

Timing changes what you learn. A request sent the day after delivery gathers first impressions, which are mostly about packaging, appearance and whether it arrived intact. A request sent after a month gathers durability and actual use, which is where the expensive complaints live. Neither is wrong, but a shop that only ever asks on day one is systematically blind to the failures that arrive in week six, and those are the failures that drive returns and refunds.

One thing not to do: prompting the customer with a suggested phrase or a draft. Beyond the legal exposure around authorship, it homogenises the corpus. Two hundred reviews containing your own marketing language teach you nothing about how buyers describe the product, which was one of the three outputs worth extracting.

What about reviews left on marketplaces and social platforms?

They count as data even when they do not count as reviews on your site. If you also sell through a marketplace, that review pile is often larger and more candid than your own, because the buyer has no relationship with you and no reason to soften anything. If those reviews arrive in several languages, translate them for analysis but publish them carefully, for the reasons set out in the piece on running a shop in more than one language.

You cannot import those reviews into your own markup. Google's documentation prohibits aggregating reviews from other websites, and doing it anyway risks the eligibility of the pages where you have legitimate product reviews. But nothing stops you reading them and running the same extraction. The aspects and broken expectations found there apply to the same product, and often surface complaints your own audience is too polite to file.

The same applies to comment threads, unboxing videos and the questions people ask before buying. Pre purchase questions are the mirror image of broken expectations: they tell you what your page failed to answer before the sale rather than after it. A shop that runs both analyses tends to find the same three gaps twice, which is a useful confirmation that the gaps are real.

What are you legally not allowed to do?

Generate reviews, incentivise a particular sentiment, or quietly delete the bad ones. All three are now explicitly named, and the first one is named as an AI problem.

The FTC's final rule on consumer reviews prohibits reviews that misrepresent that they are by someone who does not exist, such as AI-generated fake reviews. It also bans compensation or incentives conditioned on a review expressing a particular sentiment, bans undisclosed reviews by officers and managers, bans presenting a site you control as an independent review source, and bans using threats or false accusations to suppress a negative review. The Commission can seek civil penalties against knowing violators.

Read the incentive clause carefully because it is the one honest shops trip over. Offering a discount for leaving a review is not the problem. Offering a discount for leaving a positive review is, and the rule notes the condition can be conveyed implicitly. A follow up email saying you would love a five star rating in exchange for a voucher has made the condition explicit enough.

Note

Using a model to draft your reply to a review is fine. Using one to draft the review is the thing the rule names. The line is authorship of the customer's words, and it is not a grey area.

What does the EU require on top of that?

Disclosure of your method. If you display consumer reviews, EU rules require you to state how they are obtained and checked, and how you ensure they come from people who actually bought or used the product.

The obligation is more specific than most shops realise. Traders are expected to explain the verification measures used, whether all reviews are published, and how the average score is calculated, with that information on the same interface where the reviews appear or reachable by link from it. Publishing only the positive ones and deleting the rest is separately prohibited. Penalties can reach 4% of annual turnover in the member states concerned, or two million euros.

For a small shop this is a paragraph, not a project. Something like: reviews can be left by anyone who has an order reference, we publish all of them including negative ones, we remove only abuse and personal data, and the score shown is the mean of all published ratings. If that paragraph would be untrue, the fix is to the process rather than the wording.

Why does Google refuse to show your stars?

Because the reviews are about you and they are on your site. Google's review snippet documentation states that if the entity being reviewed controls the reviews about itself, those pages are ineligible for the star feature, and it says explicitly that this holds even with an embedded third party widget.

This surprises people who have paid for a widget partly to get stars in search results. The distinction Google draws is between reviews of a product, which can be eligible, and reviews of the business itself hosted by that business, which are not. Product review markup on a product page is a different case from a testimonials page about your company.

The same documentation prohibits aggregating reviews from other sites into your own markup, and prohibits fake or undisclosed incentivised reviews, which puts it in the same place as the FTC rule from a different direction. Three separate regimes, one conclusion: reviews have to be real, collected without steering, and displayed honestly.

PracticeUnited States positionEU positionGoogle position
Model written reviewsNamed as prohibited conductProhibited as fake reviewsProhibited in markup
Incentive for a positive reviewProhibited, condition may be implicitProhibited as misrepresentationProhibited or must be disclosed
Deleting negative reviewsSuppression provisions applyExplicitly prohibitedUndermines eligibility
Stars for reviews about your own businessNot a legal questionNot a legal questionIneligible for the snippet
Saying how you verify reviewsNot required by the ruleRequired on the same interfaceNot required

How do you run the analysis without leaking customer data?

Strip the identifiers before the text goes anywhere. Review bodies frequently contain names, order numbers, addresses and occasionally phone numbers, because customers write them there when they are frustrated.

A simple pass that removes anything matching an order reference, an email address or a phone number costs nothing and removes the majority of the risk. Where reviews are public anyway the exposure is lower, but the joined data is not: the moment you attach return records or purchase history to run the interesting analysis, you are handling personal data and it should be treated that way. The classification approach we set out in the policy template with data classes covers where that line sits for a small business.

One practical note on method. Do not paste four hundred reviews into a chat window and ask for themes. You will get a fluent summary that is impossible to check and that silently over weights whatever appeared near the end. Process them in batches with a fixed instruction per review, keep the per review output, and aggregate the results yourself. That way every count you act on traces back to specific reviews you can read.

How many reviews do you need before this is worth doing?

Fewer than you think for expectations, more than you think for ranking. Broken expectations show up in tens: if three people out of forty thought the item was waterproof, that is a page problem and you do not need statistics to act on it.

Ranking aspects by frequency is where small numbers mislead. With sixty reviews the difference between the second and third most common complaint is noise. Treat the top one or two as real and the tail as anecdote, and resist reorganising your roadmap around a single vivid review, which is the failure mode of doing this by hand and one that a model does not automatically cure.

What to do with what you find

Fix the page before the product, every time. A false expectation created by your own copy costs a return, a refund and often a bad review, and it is the cheapest thing in this entire article to correct.

Then decide what is worth changing about the item, using the return linkage rather than the complaint counts. Then rewrite the description in the customers' vocabulary rather than yours, which improves the page for readers and for the assistants now summarising product categories. That last point connects to what Adobe found when it scored product pages as the least machine readable page type on a retail site: the text that answers a buyer's real question is also the text a model can quote.

Finally, reply to reviews, including the bad ones, in public and briefly. That is for the next reader rather than the reviewer. A calm reply explaining that batteries are not included, under a review complaining that batteries were not included, does more for conversion than the review costs you.

"Word polarity varies in different domains, it is impossible to develop a universal opinion lexicon."Toptal, on sentiment analysis accuracy

All of this assumes you can edit the pages the analysis points at. The output of a good review pass is a list of specific copy changes: a sentence about fit under the sizing chart, a line about what is in the box, a corrected photograph caption. On a platform where copy changes queue behind a theme update, that list decays. When the storefront is yours, it is an afternoon of small edits, which is the mundane and real advantage of building a shop you can change yourself. The analysis is only worth as much as your ability to act on it quickly.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building