BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Industry/What Actually Moves Your Citation Rate in AI Answe…
IndustryAugust 10, 2026
Read · 5 min
generative engine optimization · geo

What Actually Moves Your Citation Rate in AI Answers

One study measured nine content changes against 10,000 queries. Three moved citations, one was worse than useless, and the effect size depends on your domain.

Key takeaways
  • The only peer reviewed number in this field comes from the GEO paper accepted at KDD 2024, which reports visibility gains of up to 40 percent in generative engine responses from content side changes.
  • The three techniques that moved the needle were adding quotations, adding statistics and citing sources. Keyword stuffing was the weakest of the nine tested.
  • The paper's own headline finding is the one most people drop: effect size varies by domain, so 40 percent is a ceiling observed somewhere, not a rate you should expect.
  • An answer is assembled in three stages and you can only influence the last one. Retrieval is decided by ordinary ranking, which is why Google says no special optimization exists.
  • You cannot see your citation rate, so measure it by hand. A fixed query set, a fixed cadence and four columns is enough to detect a real change.

Sites are losing clicks while their names keep turning up inside AI answers with nothing to click. That combination has produced an entire vocabulary in about eighteen months, and almost none of it is measured. The exception is a paper.

In late 2023 a group at Princeton published GEO: Generative Engine Optimization, later accepted to KDD 2024. It built a benchmark, applied nine specific content changes to real pages, and measured what happened to their visibility inside generated answers. The headline result is that these changes can raise visibility by up to 40 percent. The result that gets quoted less, and matters more, is the sentence immediately after it: the effectiveness of these strategies varies across domains, which is why domain specific optimization is needed.

Everything below sits on that distinction. There is real evidence that content changes shift citation rates. There is no evidence for a universal 40 percent.

What did the study actually test?

Nine content modifications, applied to source pages, measured against a benchmark called GEO-Bench of 10,000 queries drawn across domains and split into training, validation and test sets. Visibility was scored two ways. The first is position adjusted word count, an objective measure of how much of your page's substance survived into the answer and how prominently. The second is subjective impression, a composite of seven judgements including relevance, influence, uniqueness and perceived position.

Two metrics rather than one is the detail worth pausing on. A page can be mentioned and contribute nothing, or contribute heavily and be cited once in passing. Any tool that reports a single citation count is measuring the first and telling you it measured the second.

TechniqueWhat it changes on the pageDirection in the studyDoes it also help classic ranking?
Quotation additionAdds relevant direct quotations from named sourcesAmong the strongest gainsNeutral to mildly positive. Quotes add nothing to keyword coverage but raise perceived depth.
Statistics additionReplaces qualitative claims with quantitative onesAmong the strongest gainsPositive. Specific figures attract links, which is a classic ranking input.
Cite sourcesAttaches credible citations to claimsAmong the strongest gainsPositive, and it is already what quality guidance asks for.
Authoritative toneRewrites in a more confident, persuasive registerGains, concentrated in argumentative domainsNeutral. Tone is not a ranking input on its own.
Easy to understandSimplifies vocabulary and sentence structureModestPositive for engagement, neutral for retrieval.
Fluency optimizationSmooths the prose without adding informationModestNeutral. This is the classic empty rewrite.
Unique wordsIntroduces rarer vocabularyMinimalNeutral to negative if it costs clarity.
Technical termsAdds specialised terminologyMixed, domain dependentPositive only where the audience uses those terms.
Keyword stuffingRepeats query termsWeakest of the nineNegative. It has been a spam signal for two decades.

Read the last column, because almost nobody separates these two questions and it is where the practical decisions live. Three of the nine techniques are things a careful editor already does. Two are neutral polish. One is actively harmful in both systems at once. The overlap between what helps an AI answer and what helped a search engine ten years ago is much larger than the discourse suggests, and the divergence is narrower and more specific than the word optimization implies.

One further finding deserves its own line. The gain was not evenly distributed by starting position. Lower ranked sources benefited most, with a rank five source reported as gaining up to 115 percent visibility from citing sources. If you are already the top result, these techniques defend a position you hold. If you are on the second page, they are one of the few levers that changes the shape of the outcome rather than the margin.

Which technique wins depends on the kind of question

The paper's domain breakdown is the part that should change how you read every generic recommendation you have been given. The authoritative rewrite performed best on debate and history questions, where the answer is an argument and confidence reads as substance. Statistics addition performed best on law, government and opinion questions, where a claim without a number is indistinguishable from an assertion. Quotation addition performed best on people and society questions and again on history, where the object of interest is what a specific person or document said.

Line those up and a pattern falls out that the paper does not spell out. Each winning technique supplies the evidence type that its domain treats as proof. A history question wants a source. A policy question wants a figure. An argumentative question wants a position stated clearly enough to quote. That reframes the whole exercise: you are not applying nine techniques, you are working out what counts as evidence in your subject and then supplying more of it than the pages around you do.

For a commerce or product question, which is where most of this blog's readers operate, the equivalent evidence type is usually a specific measured outcome with the conditions attached. Not a claim that a technique improves conversion, but the number, the sample, the period and the segment. That is a harder sentence to write and it is the one that gets cited. It is also what the largest retailers landed on in 2026, as the brands rebuilding product pages for AI search found: stated benefits, product level FAQs and usage detail beat any tool purchase.

How does an engine decide what to cite?

In three stages, and understanding which one you can touch saves an enormous amount of wasted effort.

Sequence diagram showing retrieval then synthesis then citation, marking the citation stage as the only one a publisher can influence directly

Retrieval selects candidate documents. This is ordinary ranking. Google is unusually direct about this in its guide to optimizing for generative AI features, which states that the best practices for search engine optimization continue to be relevant because the generative features are rooted in the core ranking and quality systems. The same document tells you that structured data is not required for generative search, that there is no special schema, and that breaking content into tiny pieces is not needed because their systems handle multi topic pages.

Synthesis writes a single answer from several documents at once. Nothing you wrote reaches the reader intact at this stage. Your sentences are compressed, merged with a competitor's and rephrased. This is the stage that most advice implicitly targets and the stage you have least influence over.

Citation attaches links to the claims the model leaned on. This is the stage the GEO paper measured, and it explains why its three winning techniques are what they are. A direct quotation is an object that survives compression because it cannot be paraphrased without losing its point. A statistic is an atom of information the model cannot generate on its own and must attribute. A cited source signals that the claim came from somewhere checkable. All three make a passage easier to lift and harder to lift anonymously.

Note

This is why the same page can gain AI citations and lose clicks at the same time. Making a passage more quotable makes the answer more complete, which reduces the reason to visit. Citation share and click share are not the same objective, and the honest version of this work admits they sometimes pull against each other.

Is GEO different from AEO, or from SEO?

The first pair, no. The Wikipedia entry on generative engine optimization states plainly that as of early 2026 there is no consensus in the academic literature distinguishing GEO, answer engine optimization, artificial intelligence optimization and LLMO, and that practitioners use them interchangeably. If a vendor claims a methodological difference between two of these acronyms, they are selling a distinction the field has not made.

The second question is more interesting and the answer is a qualified partly. Google's position, quoted in the same entry, is that optimizing for generative AI search is optimizing for the search experience and is therefore still SEO. Forrester's Nikhil Lai has argued that the terms are significantly but not fundamentally different. Both readings are compatible with the study: the retrieval stage really is classic ranking, and the citation stage really does respond to changes that classic ranking is indifferent to.

The practical translation is a priority order rather than a new discipline. If you do not rank, none of this applies, because you never enter the candidate set. If you rank and are not cited, the citation stage is where to work.

Why can you not see your own citation rate?

Because the reporting is aggregate, sampled and split across vendors, and it will stay that way. Google's generative AI performance report in Search Console covers AI Overviews and AI Mode, gives impressions with breakdowns by page, country, device and date, and comes with a set of limitations worth reading before you build anything on it: it rolled out to a subset of owners, it needs sufficient impressions to show anything at all, it excludes Search Labs experiments, and it carries the same 1,000 row cap as the standard report. It does not let you separate one AI surface from another.

Microsoft ships a comparable report inside Bing Webmaster Tools covering Copilot citations and the grounding queries the assistant generates internally. Its limitations run the same way: grounding queries are sampled rather than exhaustive, surfaces are aggregated, and the metric is citation frequency rather than prominence within the answer.

Then there is everything neither of them sees. Assistants that do not publish publisher tooling, the conversational surfaces inside other products, and the share links that leak into the index by accident, which we looked at when shared chat transcripts turned up in Google's index. The measurement gap is not a temporary tooling problem. It is structural, because the engines have no commercial reason to close it.

A sampling protocol you can actually run

Manual sampling is the only method available to a small team, and it is more useful than it sounds, because you are not trying to measure an absolute rate. You are trying to detect whether a change you made moved anything. That is a much easier measurement and it needs only consistency.

Build the query set once

Twenty to forty questions, written as a person would type them, not as keywords. Split them deliberately: about half should be questions where you already rank in the top five, because those test the citation stage in isolation, and about half should be questions you want to win, because those test whether you enter the candidate set at all. Freeze the list. A query set you edit is a query set that cannot show a trend.

Fix the cadence and the conditions

Monthly is enough. Weekly produces noise you will misread as signal. Run every query in a logged out session, in the same country, on the same set of engines, in the same order, on the same day of the month. Model updates and index refreshes will still move your numbers, which is exactly why the conditions you control have to stay still.

Record four columns and nothing else

ColumnWhat you write downWhy this one
CitedYes or no, per engineThe base rate. Everything else is meaningless without it.
Position in the answerWhich paragraph carried your link, first, middle or lastA citation in the closing line is worth much less than one supporting the main claim.
Claim usedThe specific sentence or figure of yours that the answer leaned onThis is the column that tells you what to write more of. It is also the only one that generalises.
Who else was citedThe other two or three domains in the answerYour real competitive set for this question, which is often not the set ranking above you.

After three months you will have something no tool can sell you: a list, in your own words, of which of your claims get borrowed. If the study generalises to your domain, that list will be dominated by passages carrying a specific number or a named source, and the claim column is how you find out whether it does. That is the whole point of running the sample rather than trusting the headline figure.

Card listing the four columns to record in a manual citation sampling run across generative answer engines each month

What to change on the page, in order

The study gives an unusually clear priority list, and it is short.

  1. Replace one vague claim per section with a sourced number. Not a rounded market size. A specific figure with a date, a method and a link to where it came from. This is the single highest leverage change and it also survives the next algorithm update, because it is just better writing.
  2. Add a real quotation where you currently paraphrase. A named person or document saying a thing in their own words is an object that resists compression.
  3. Attach the source to the claim, inline, not in a list at the bottom. A footer of references is invisible to a system reading a passage in isolation.
  4. Answer discrete questions under their own headings. A passage that is self contained is a passage that can be lifted whole.
  5. Delete the keyword repetition. It was the worst performer of the nine and it has been a spam signal in classic search for twenty years.

What is absent from that list is as informative as what is on it. No special file, no new schema, no chunking. Google states directly that files of that kind are neither harmful nor helpful for its own search, which matches what we found when we went through the evidence on whether llms.txt does anything: adoption is real, measured effect is not.

You rank and you are not cited. What now?

This is the most common situation and it has three distinct causes that need different responses, so the first job is to tell them apart using the claim column from your sample.

Your page has nothing liftable. The answer covers your topic, cites two other domains, and neither of their claims is one you made. Look at what got cited: if the other pages carry figures and yours carries adjectives, you have the answer and the fix is the priority list below. This is the cause in the majority of cases and it is the cheapest to correct.

Your page is liftable but buried. The claim column shows one of your sentences being used, but the citation lands in the closing paragraph of the answer while another domain carries the main point. You are in the candidate set and you are being treated as supporting evidence. The correction is structural rather than editorial: the claim that should be central is probably three screens down under a heading that does not ask the question the reader asked.

The question does not want a source. Some answers cite nothing, because the model considers the material common knowledge. No amount of editing changes that, and continuing to work on those queries is the most reliable way to waste a quarter. Mark them in the query set and stop counting them against yourself.

The reason for insisting on a written claim column rather than a citation tally is precisely this: a yes or no count cannot distinguish between these three, and every one of them looks identical on a dashboard.

What is being sold that the evidence does not support?

Three things, and each is worth recognising by shape rather than by vendor.

The first is a special file or markup that makes you legible to AI. Google's own guidance says its search does not use files of that kind and that they neither help nor harm, and it says there is no special schema for generative features. A vendor is free to argue that other engines behave differently, but that argument needs its own evidence and is almost never offered with any.

The second is a citation share number presented without a method. If a tool reports that you hold some percentage of citations in your category, ask what query set produced it, how often it runs, and whether it is logged out. Absent those, the number is not comparable to anything, including its own value last month.

The third is the promise of the 40 percent. It is a real number from a real paper, and it is an upper bound observed across a benchmark, not a forecast for your site. The paper says in its own abstract that efficacy varies across domains and that this is why domain specific methods are needed. Quoting the ceiling while omitting the caveat inverts the study's actual conclusion.

Does any of this depend on the content being human written?

Indirectly, and the mechanism is worth naming precisely. None of the three winning techniques can be executed by a generator working from its own knowledge. A quotation requires a real source. A statistic requires a real measurement. A citation requires a real document at the other end. Text produced without leaving the model has none of these, which is why it tends to score badly on exactly the axes the study found to matter.

That is the useful reframing for anybody producing content at volume. The question is not whether a machine helped write the sentence. It is whether the page contains anything that had to come from outside a model, which is also close to how detection systems have started to reason, as we covered in what changed in AI content detection. A page assembled entirely from what a model already knows competes with the model directly, and loses, because the answer engine can produce that content itself without citing anyone.

If you are exposing structured data to assistants rather than only publishing prose, the same principle applies at the interface level: what you expose has to be something the model cannot infer. That is the reasoning behind how we expose store data over MCP, and it is the same bet as the one above, made in a different format.

What to expect, honestly

Between two and six months before a manual sample shows anything, because you are measuring a slow moving system with a small sample. A change in citation rate that does not also change your ranking is the clearest signal you will get that the work is doing something specific rather than riding a general improvement.

And a warning that follows from the study's own caveat. Because effect size varies by domain, the first thing to establish is not how much lift the techniques give in general. It is how much they give in yours. That is what the fixed query set is for, and it is why the protocol above starts with measurement rather than with edits. Anyone quoting 40 percent at you as a forecast has read the abstract and stopped.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building