The best measurement of what online reputation work is worth comes from tens of thousands of hotels, and the headline number is smaller than any agency will quote you. Ratings rose by an average of 0.12 stars. Not half a star. About a tenth of one.
The interesting part is not the size of the number. It is the mechanism that produced it, because that mechanism is the thing a badly deployed AI removes.
- Hotels that began responding to reviews saw ratings rise by an average of 0.12 stars and received 12 percent more reviews, measured across tens of thousands of TripAdvisor listings.
- The same study found negative reviews became less frequent but longer, which points at the real mechanism: people write differently once they know somebody is reading.
- That mechanism is the asset, and a reply that reads as machine generated is the one way to spend the effort and lose the benefit at the same time.
- The FTC's position is short and quotable. You may respond publicly, and you should watch what you say. False accusations against a reviewer and groundless legal threats are prohibited.
- Google screens replies against its content policies before publishing, usually in ten minutes but sometimes taking up to thirty days, so a bulk run can fail silently.
- The dangerous failure is not a clumsy sentence. It is an assistant handed the order record that publishes a customer's name, address or order number in a public reply.
Does replying to reviews actually change anything?
Yes, measurably, though by less than the marketing suggests. The best available evidence is the study by Davide Proserpio of USC Marshall and Georgios Zervas of Boston University's Questrom School, published in Marketing Science and built on tens of thousands of hotel reviews and management responses from TripAdvisor.
Their headline findings, as summarised in the INFORMS release on the research, are two numbers. Hotels that started responding saw ratings increase by an average of 0.12 stars, and they received 12 percent more reviews than comparable hotels that did not respond.
A tenth of a star sounds like nothing until you consider how ratings behave. Aggregate scores move very slowly, they cluster tightly between competitors, and the platforms rank on them. On a listing page where every business sits between 4.1 and 4.5, a tenth of a star is not noise. It is placement.
The 12 percent volume increase matters more than the rating for a small shop, because volume is what makes a rating credible. A 4.6 across nineteen reviews and a 4.6 across two hundred are not the same signal to a buyer, and responding is one of the few levers that moves the denominator without touching the FTC's rules on incentives.
Why does an obviously automated reply undo the benefit?
Because the benefit was never the text. It was the evidence that somebody read the complaint, and a generic reply is evidence of the opposite.
Look again at the study's second finding, which is the one that almost never gets quoted. When hotels began responding, negative reviews became less frequent and simultaneously longer. That is a strange pairing until you work out what causes it. People who were about to fire off a short, indefensible complaint stopped bothering, because someone was clearly going to read it and answer. People with a genuine grievance wrote more, for the same reason.
The reply, in other words, functions as a signal of scrutiny. It tells the next unhappy customer that this page is watched. That is what suppresses the throwaway one star and what draws out the detailed complaint you can actually act on.
Now imagine that same page filled with forty replies that all begin "We're so sorry to hear about your experience" and all end with an invitation to contact the team. The signal inverts. A reader can tell in two seconds that nobody read anything, and the deterrent effect that produced the measured gain disappears. You have paid for the replies and bought the opposite of what they were for.
This is the specific reason the usual advice to automate review replies is backwards. The task looks like a perfect candidate for automation: repetitive, high volume, low creativity. It is in fact a task whose entire value comes from being visibly non automated, which is an unusual property and the reason it deserves a different rule than the rest of your queue.
So what should the assistant actually do?
Draft, retrieve and check. Not send. An AI review response pipeline is worth building, as long as the last step stays manual. The split below is the one that keeps the measured benefit while removing most of the work, and it is organised by what the reply has to prove rather than by review length or star rating.
| Review type | Safe to draft with AI | What a human must add | Never automate |
|---|---|---|---|
| Positive, no detail | Yes, and sending is low risk | Nothing, or one specific word from the review | Nothing |
| Negative, factual complaint | Yes, as a draft with the order facts retrieved | The specific fact and what you did about it | Sending unread |
| Negative, disputed facts | Draft only, no claims | The entire substance | Any statement about what the customer did |
| Suspected fake or malicious | No | All of it, or a platform report instead | Any accusation, any legal reference |
The second row is where nearly all the value sits. An assistant that pulls the order date, the item, the delivery record and your policy into a draft has done ninety percent of the labour, and the remaining ten percent is the sentence only you can write, which is what was actually wrong and what you changed. That sentence is also the one a reader recognises as human, so it is doing double duty.
Reading the reviews at scale is a separate job with a separate tool, and one where automation is unambiguously the right call. We covered what that reading finds and misses in our piece on what AI actually extracts from customer reviews.
What does a reply that reads as human actually contain?
One fact the reviewer supplied, one thing you did, and no apology template. Everything else is optional, and most published replies get the proportions exactly backwards by leading with sympathy and never arriving at substance.
Compare two replies to the same complaint, which said a chair arrived with a scratch on the left arm and took nine days rather than the three that were promised.
The first: "We're so sorry to hear about your experience. This is not the standard we hold ourselves to. Please reach out to our team so we can make this right." It is polite, it is grammatical, and it could be attached to any negative review about any product on earth. A reader learns nothing, and the next unhappy customer learns that nobody is really reading.
The second: "Nine days for a three day delivery is our error, not the carrier's. We switched the chair line to a different depot in July and the handover has been adding a week. It is fixed as of last Tuesday. The scratch on the arm should not have passed our check and we are replacing it." A reader learns that the business knows what went wrong, that it was systemic, that it has been corrected. The next reviewer learns that the page is watched.
Notice what the second reply does not contain. No name, no order number, no address, no claim about what the customer did. It is specific about your operation and vague about their identity, which is the correct direction for a public page and the opposite of what an assistant with database access will produce unprompted.
Notice also that the second reply is not longer. Length is not the signal and padding a template does not fix it. The difference is that one sentence in the second reply could not have been written without reading the review, and one could not have been written by anyone outside the business.
Where does this leave a shop with two hundred unanswered reviews?
Answering the recent ones and leaving the rest, which is the opposite of what most catch up projects attempt. A backlog run is where automated replies do the most damage, because two hundred replies posted in one afternoon are visibly a batch no matter how well each one reads.
The deterrent effect the study measured operates on future reviewers, not past ones. A reply posted eighteen months after the review reaches almost nobody: the reviewer has moved on, and a reader scanning your page discounts it as housekeeping. The value of a reply decays fast, which means the backlog is worth much less than the flow.
The practical order is the recent negatives first, because they are the ones prospective customers actually read and the ones still capable of being resolved. Then the recent positives, which are cheap and where automation is genuinely fine. Then stop. The reviews from two years ago are a sunk cost and answering them buys nothing you can measure.
One exception is worth making. If an old negative review contains a factual claim that is no longer true, because you changed the policy or fixed the fault it describes, a reply saying so is worth posting whatever its age. That reply is not reputation management, it is a correction, and it is the one case where a reader genuinely benefits from an answer to an old complaint.
For anything beyond that, the effort is better spent making the flow of new reviews larger. The same study found responding lifted review volume by 12 percent on its own, and the tools that ask well at the right moment do more for a thin review count than any amount of retrospective replying. Just keep the asking clear of the incentive rules, since conditioning anything on a positive review is exactly what the FTC rule prohibits.
What are you not allowed to say?
Less than you think, and the boundaries are published. Two authorities matter for most sellers and neither is difficult to comply with once you have read them.
The FTC's guidance on its consumer reviews rule answers the question directly. Asked whether a business can respond to a negative review, the Commission's own questions and answers page gives about the shortest useful answer in all of consumer regulation.
What it prohibits is review suppression, and the definition is broader than most sellers assume. Making a false accusation about a reviewer, knowing it is false or with reckless disregard for whether it is, falls inside the rule. So does intimidation, which the FTC describes as extending well past physical threats to abusive communications and character assassination used to deter someone from acting. Unfounded legal threats are prohibited; a legitimate legal action with a proper factual basis is not.
Read that alongside the fourth row of the table above and the reason for the never automate column becomes obvious. An assistant asked to reply firmly to one of the fake reviews you believe you have received will happily produce an accusation, because that is what firmly means. The rule it just broke is not one it knows about. Our piece on what regulators actually ban around reviews covers the wider rule, and the suppression provisions are the half that catches honest businesses.
The second authority is the platform. Google Business Profile guidance on managing customer reviews states that it reviews replies against its content policies before they publish, and that an unapproved reply must be edited before it appears. Replies usually clear in about ten minutes, but Google says a review can take up to thirty days.
The trap that costs more than a bad sentence
An assistant with access to your order records will use them, and a public reply is the wrong place for most of what it finds. This is the failure that turns an efficiency project into a personal data incident, and it happens because the integration that makes drafts good is the same one that makes them dangerous.
The useful draft needs context. It needs to know the customer ordered on the fourteenth, that the parcel was delayed by the carrier, and that your policy allows a replacement. So you connect the assistant to the order system, and the drafts improve immediately. The problem is that nothing in the pipeline knows which of those facts is safe to publish under your shop name on a public page.
A reply that opens with "Hi Sarah, I can see order 48219 was delivered to the Oakfield Road address on the sixteenth" is helpful, accurate, and publishes a customer's name, order number and street in a place anyone can read. It also runs directly into Google's prohibition on posting private or confidential information, including contact information linked to a name.
The fix is a rule at the output rather than a smaller prompt. Nothing generated for a public reply may contain an order identifier, a surname, an address, a phone number, an email, or any fact the reviewer did not themselves state in the review. That last clause is the one that catches the subtle cases, and it is easy to check mechanically before anything is sent.
Everything the draft retrieved is still useful. It belongs in the private message or the email that follows, where the customer is identified and the channel is not public. Splitting the two outputs, one public and vague, one private and specific, gets you the good draft without the exposure.
How do you tell if your replies are working?
By watching the shape of the reviews rather than the rating, because the rating moves too slowly to tell you anything for months. The study gives you the two indicators to track, and both show up long before a tenth of a star does.
The first is the frequency of short negative reviews. If the mechanism is working, the one line one star with no detail should get rarer, because those are the reviews the deterrent effect removes. Count them monthly. This is the earliest signal available and it moves within weeks.
The second is the length of the negative reviews you do get. Longer is the expected direction and it is a good sign, uncomfortable as it feels. A detailed complaint is a customer explaining what to fix, and the same study found that is exactly what responding produces.
The third indicator is not in the study and is worth adding: the proportion of your replies that contain a fact specific to that review. Sample twenty of your published replies and count how many mention something only that customer said. If the number is low, you have automated the form and lost the function, whatever the volume looks like.
Cost is worth measuring too, because drafting at volume is a running expense rather than a one off. Anyone weighing that against the manual alternative should look at what generating this kind of content actually costs per month before committing a queue to it.
The rule that survives every platform change
Write a review reply that could only have been written by someone who read the review. That single test resolves nearly every question in this article without needing to remember which platform screens what.
It tells you to draft with an assistant and finish by hand. It tells you why the fourth row of that table is never automated. It tells you which facts belong in public and which belong in an email. And it happens to be the same standard the measured benefit rests on, which is not a coincidence: the 0.12 stars were never paid for by the words. They were paid for by the reading.