- Across seven datasets and more than 880,000 texts, language models cut writing complexity variance by 21% to 50% while leaving the substance intact.
- The same research found that polishing a text strips signals of who wrote it, including cues to age, ideology and moral outlook.
- Most owners try to fix generic copy by asking for a friendlier tone. Tested on 100 people, trustworthiness explained about 52% of whether a brand was desirable and friendliness added only about 8% more.
- A playful tone for an insurance brand raised friendliness scores, lowered trustworthiness and did nothing for willingness to recommend.
- A usable voice brief is mostly a list of words you refuse to use and a handful of sentences you actually wrote, not a set of adjectives.
- Google Merchant Center requires AI generated product images to carry IPTC metadata marking them as trained algorithmic media, and product data to be marked separately.
Ask a model to write a product description and it will hand you something competent in four seconds. Read twenty of them and you notice the problem: they are all the same person. Not badly written, just written by nobody in particular, in a register that sits somewhere between a press release and a friendly email.
This is not a failure of prompting and it is not fixed by asking for more personality. It is a measurable property of what these tools do to text, and once you know its shape you can work around it.
What does a language model actually do to writing?
It keeps what you said and flattens how you said it. The content survives editing. The individuality does not.
A study published in Nature Human Behaviour and available as a preprint on the shrinking landscape of linguistic diversity ran three studies across seven datasets and more than 880,000 texts. Its central finding is narrow and useful: models reduce writing complexity variance by a statistically significant 21% to 50% across datasets and models, while preserving the substance of what was written. They amplify dominant patterns and suppress the rest.
The second finding is the one that should interest anyone with a brand. Polishing a text strips cues to gender, age, ideology and moral values. Those cues are not decoration. They are how a reader forms a sense of who they are buying from, which for a small shop is frequently the entire competitive advantage over a large one.
The researchers also observed the effect outside the lab. After ChatGPT's release, variance in writing complexity dropped across social media, news and scientific writing. The convergence is not something happening to your drafts in isolation. It is happening to everyone's, in the same direction, which is precisely why the output reads as familiar.
Is a friendlier voice a better voice?
Usually not, and this is where most attempts to fix generic copy go wrong. Friendliness is easy to generate and it is not what makes people buy.
Nielsen Norman Group tested this directly. Its study on how tone of voice changes brand perception surveyed 100 American adults after qualitative testing, presenting near identical website content that varied only in tone, across auto insurance, banking, home security and healthcare.
The numbers are worth memorising. Trustworthiness alone explained roughly 52% of the variability in whether people found a brand desirable. Friendliness added only about 8% beyond that. And a playful tone for the insurance brand increased how friendly it seemed while reducing how trustworthy it seemed, with no improvement in willingness to recommend.
So the instruction most people give a model, make it warmer and more fun, optimises the dial that barely moves the outcome and can damage the one that does. Casual is not the enemy here. The bank tested with a casual tone beat its formal version on both trustworthiness and recommendation. Jokes are the risk, particularly in categories where the reader is worried about something.
What are you actually adjusting when you describe a voice?
Four things, and naming them is more productive than reaching for adjectives. Nielsen Norman Group's framework breaks tone into four spectrums that you can position yourself on deliberately.
The four dimensions of tone of voice are formal against casual, serious against funny, respectful against irreverent, and matter of fact against enthusiastic. Each is a spectrum rather than a switch, and almost no effective business writing sits at an extreme of any of them.
| Dimension | What it changes for a shop | Where the evidence points | What to tell a model |
|---|---|---|---|
| Formal or casual | How close the reader feels to the person selling | Casual raised both trust and recommendation for the bank tested | A position, not a word: closer to casual, short sentences, no slang |
| Serious or funny | Whether a reader relaxes or gets nervous | Humour risked alienating readers when it missed or obscured information | Default to serious in anything about money, safety or delays |
| Respectful or irreverent | Who the joke is aimed at | Irreverence works against the subject, not the audience | Irreverent about the category if you like, never about the customer |
| Matter of fact or enthusiastic | How much the copy insists you should be excited | Healthcare readers found enthusiasm reassuring under stress | Enthusiasm about specifics, never about the product in general |
The last row describes the most common failure in generated product copy. A model asked to be enthusiastic produces enthusiasm about the object as a whole, which reads as advertising. A person who makes the thing is enthusiastic about one detail, the stitching or the way the lid seals, which reads as knowledge. The difference is not tone at all. It is whether the writer knows anything.
How do you write a brief that actually carries a voice?
With examples and prohibitions rather than adjectives. Telling a model to be warm, authentic and conversational produces the average of every text ever described that way, which is exactly the register you were trying to escape.
Three components do most of the work.
A list of words you never use. This is the single highest leverage part of a voice brief and almost nobody writes one. Nielsen Norman Group calls these anti tone words. Yours might include elevate, curated, effortless, indulge, must have, game changer, or whatever the generic version of your category reaches for. A prohibition is checkable in a way that an aspiration is not: you can search a draft for a banned word, you cannot search it for authenticity.
Three to five sentences you genuinely wrote. Not a style description, the actual sentences, ideally from an email to a customer rather than from your website, because website copy has usually already been sanded down. These give the model something to imitate that has the cues the research says get stripped. Include one sentence that is slightly awkward. The awkwardness is the fingerprint.
The reader, described as a person. Not a demographic. Somebody who has already read three other shops' pages this evening and is trying to work out whether yours is a real business. What that person needs is almost never more enthusiasm, which brings the brief back in line with the trustworthiness finding above.
Written down, that brief is half a page and it does more than any amount of prompt engineering. We went through the structural side of the same problem in a piece on a method for writing product descriptions with AI, where the argument is that the model should be given facts and denied claims.
What should never be generated at all?
Anything that is a promise, and anything that is the reason somebody chooses you over a cheaper option. The first is a liability question and the second is a strategy one.
Promises are the obvious category: delivery windows, returns terms, guarantees, claims about materials or provenance. A generated sentence stating a policy is a commitment your business may be held to, which is the ground covered in our piece on what happens when a shop's automated words become a promise.
The second category is subtler and costs more over time. If people buy from you because you know something, the sentences where that knowledge shows are the ones that must be yours. For a maker that is the paragraph about why a material was chosen. For a reseller it is the honest note about which of two similar products is actually better for which use. A model can write around that knowledge fluently enough that you will not notice it has been replaced with plausible filler until the page stops converting.
The practical split is boring and effective. The model drafts structure, variations and the parts of a catalogue that are genuinely repetitive. You write the one paragraph per product that could only have come from someone who handled the thing. At two hundred products that is still a lot of writing, and it is the writing that pays.
Does Google care whether a person wrote it?
Not directly, and the distinction it does draw matters more than the one people expect. Google's position is about value and scale rather than authorship.
Its guidance on generative AI content says using AI to generate many pages without adding value for users violates the scaled content abuse policy, while acknowledging generative AI can be valuable for research and for structuring original content. The instruction is to focus on accuracy, quality and relevance, with explicit attention to generated titles, meta descriptions, structured data and image alt text, because those surface in search results.
There is one requirement in that page that surprises most merchants. Google Merchant Center asks for AI generated images to carry IPTC metadata marking them as trained algorithmic media, and for product data to be marked as AI generated separately. That is a labelling obligation attached to your product feed rather than an abstract quality principle, and it applies whether or not anyone would notice the image was generated.
Google also suggests giving readers background on how automation was used where it makes sense. That is a lighter touch than a disclosure rule, and for a small shop the honest version tends to help rather than hurt, because it is exactly the kind of detail a generic competitor will not bother with. How generated content performs in search more broadly is a separate question we looked at in what actually ranks when the content was AI generated.
How do you tell whether any of this worked?
Show the copy to somebody who knows your shop and ask them to guess who wrote it. That test costs nothing and is more reliable than rereading your own drafts, because you cannot hear your own voice properly on a page you just approved.
Nielsen Norman Group's own recommendation is more structured: run a product reaction test, or ask people to place your content on each of the four dimensions, and check whether where they put it matches where you intended it to be. For a business with a mailing list this is one email to a dozen customers and it will tell you more than any analytics dashboard.
There is a cheaper internal check worth running monthly. Take three recent pieces of your copy and three from a competitor in the same category, strip the brand names, and see whether you can tell them apart. If you cannot, the convergence the research describes has already happened to your catalogue, and the fix is not a better prompt. It is putting some of your own sentences back.
Should the voice change from page to page?
The intensity should, the identity should not. Nielsen Norman Group's own guidance is to keep the brand consistent while adjusting how strongly the tone is applied depending on the topic and the reader's emotional state.
This is the part a model will get wrong unless you tell it. Handed one voice brief, it applies the same dial settings to a product page, a delivery delay notice and a refund refusal, and the result on the last two is jarring. A shop that is cheerfully casual about a candle and cheerfully casual about a lost parcel reads as a business that is not listening.
The healthcare result in the same study is instructive because it runs against the intuition. Participants there unanimously preferred a casual, enthusiastic approach to a formal one, finding it reassuring during a stressful moment. So the rule is not that serious situations demand formality. It is that the tone has to acknowledge what the reader is feeling, and formality is only sometimes the way to do that. Warmth during a problem is fine. Jokes during a problem are not, which is the same boundary the humour research points at from the other direction.
In practice this means your brief needs a second short section: how the voice shifts when something has gone wrong. Two or three sentences covering a delay, a refund and a refusal will cover most of what a small shop ever has to send, and those are the messages where generic copy does the most damage. Getting that right matters more than the product pages, because a customer forms their strongest opinion of a business at the moment it disappoints them. The same logic drives how we suggested handling replies to negative reviews with AI, where the temptation to sound upbeat is strongest and most costly.
How many tone words should a brief contain?
A handful, and the count is a real constraint rather than a stylistic preference. A brief listing twelve qualities describes nothing, because no piece of writing can be twelve things and the model will average them into the default register you were trying to avoid.
Pick two or three positive words and, more importantly, the same number of anti tone words. Then test whether they are doing any work by writing one sentence that obeys them and one that violates them. If you cannot easily write the violating version, the word was not specific enough to constrain anything. Approachable fails this test. Never uses exclamation marks passes it.
Where the voice actually lives
Not in adjectives, and not in the model. It lives in specifics: the things you know that a general purpose writer could not, expressed in sentences that carry the marks of a particular person having written them.
The research this article rests on is oddly encouraging on that point. What models flatten is variance, and variance is cheap for you to supply because you already have it. You have opinions about your products, a way of explaining things to customers that you have refined over years of being asked, and probably a handful of turns of phrase you use without noticing. None of that requires effort to produce. It requires only that you stop deleting it in favour of something that sounds more professional, which is usually what generic means.
If you are setting up a shop and want the copy fields, metadata and product structure under your own control rather than locked inside a theme, our ecommerce website builder generates the storefront as code you own, which makes the difference between editing your own words and negotiating with a template.
The half page brief is the whole method. A list of banned words, a few sentences you actually wrote, and a clear picture of the person reading. Then let the model do the repetitive part, and keep the paragraph that only you could have written.