Quick Answer: The GEO paper (Aggarwal et al., KDD 2024) tested nine content edits across 10,000 queries. Adding statistics, quotations and source citations lifted answer visibility 30 to 40 percent; keyword stuffing scored below the unoptimized baseline. The famous 40% measures share of the generated answer, not traffic.

Nearly every agency pitch deck that mentions generative engine optimization cites one number: up to 40 percent more visibility, from a Princeton study. Very few of the people repeating it have read the twelve pages behind it. That gap matters, because the paper is more useful than the soundbite, and also more modest. It names the tactics that worked, quantifies the ones that failed, warns that results differ by domain, and states its own expiry conditions. This teardown walks through the actual experiment: how the benchmark was built, what the metrics measure, which of the nine methods earned their reputation, and where the paper stops and the marketing starts.

The paper, its authors, and why it carries weight

"GEO: Generative Engine Optimization" was written by Pranjal Aggarwal (IIT Delhi), Vishvak Murahari, Karthik Narasimhan and Ameet Deshpande (Princeton University), Tanmay Rajpurohit and Ashwin Kalyan, posted to arXiv in November 2023 as arXiv:2311.09735, and presented at KDD 2024 in Barcelona, one of the main peer-reviewed venues for data mining research. It did three things at once: it formalized the idea of a generative engine, systems like BingChat and Perplexity that retrieve web sources and write a synthesized, cited answer; it proposed visibility metrics suited to that format; and it ran the first controlled experiment on whether content edits change how prominently a source appears in generated answers.

The framing is creator-centric. The authors' stated concern is that generative engines answer the user directly, which can cut organic traffic to the sites that supplied the information, and that small creators have no visibility into how these black-box systems select and portray their content. GEO, as they define it, is a black-box optimization framework: you cannot see the engine's internals, so you test which changes to your own pages improve your measured share of the answers.

GEO-bench: what the 10,000 queries actually are

The benchmark deserves more attention than it gets, because the results are only as representative as the queries behind them. GEO-bench contains 10,000 queries drawn from nine sources and split 8,000/1,000/1,000 for training, validation and test. Three of the sources are the standard search-research datasets built from real anonymized user queries on Google and Bing: MS MARCO, ORCAS-I and Natural Questions. The rest were chosen to stress the synthesis behaviour that separates a generative engine from a ten-blue-links page: essay questions from All Souls College, Oxford; reasoning-heavy questions from LIMA; debate questions from an earlier generative-engine evaluation; trending queries from Perplexity's Discover feed; layman-explanation questions from the ELI5 subreddit; and a set of GPT-4-generated queries added to widen domain and intent coverage.

Each query was paired with the cleaned text of the top five Google results, and each was tagged by domain and intent, with the distribution held at roughly 80 percent informational, 10 percent transactional and 10 percent navigational. The test engine itself is a two-step research build that mirrors commercial designs of the time: fetch the top five sources, then have gpt-3.5-turbo write a cited answer from them, sampling five responses at temperature 0.7 and averaging over five random seeds to damp statistical noise. Two details worth registering before you generalize: the whole apparatus retrieves only five sources per query, and the answer-writing model is a 2023-era LLM. Both choices were reasonable and both date the findings.

What "visibility" means here, and what the 40% is made of

This is the section the industry skipped. In classic SEO, visibility is your average rank across queries. A generated answer has no rank list; your site appears as inline citations of varying length, position and prominence. So the authors built two metrics.

Position-Adjusted Word Count is the objective one. It starts from the share of the answer's words that sit in sentences citing your source, then discounts by an exponentially decaying function of citation position, on the logic (supported by click-through studies) that earlier material gets read more. A source cited early and at length scores high; a source name-checked once at the bottom scores low. All citation scores in a response are normalized so they sum to one, which makes this a zero-sum share-of-answer measure: your gain is another source's loss.

Subjective Impression is the judged one. Seven facets, including relevance to the query, how much the answer relies on the citation, uniqueness of the cited material, subjective position and estimated click likelihood, are each scored by GPT-3.5 using a G-Eval-style rubric, then normalized against the objective metric so the two can be compared. An LLM grading an LLM is a legitimate methodology with known calibration limits, and the authors treat it accordingly.

Now the number everyone quotes. "Up to 40%" is the relative improvement of the best methods over an unoptimized baseline on Position-Adjusted Word Count, inside this testbed. Per Table 1 of the paper, the top methods improve on baseline by 41 percent on Position-Adjusted Word Count and 28 percent on Subjective Impression. Neither metric counts a click, a visit, a lead or a dollar. What the study demonstrates is that specific content edits reliably enlarge your slice of the generated answer. That is genuinely valuable, and it is a different claim from "GEO increases traffic 40 percent," which the paper never makes.

The nine methods on the bench

Every method is a targeted rewrite of one source website, applied by prompting an LLM to make a specific class of change. For each query, one of the five retrieved sources was picked at random and each method applied to it separately, so every tactic faced the same competitive context. The nine:

  • Authoritative: rewrite in a more persuasive, confident style
  • Statistics Addition: replace qualitative claims with quantitative ones where possible
  • Keyword Stuffing: add more query keywords, classic-SEO style
  • Cite Sources: add citations to credible sources
  • Quotation Addition: add quotations from credible sources
  • Easy-to-Understand: simplify the language
  • Fluency Optimization: improve the flow of the text
  • Unique Words: add distinctive vocabulary
  • Technical Terms: add domain terminology

Note the design intelligence: six of the nine only re-present what the page already says, while Cite Sources, Quotation Addition and Statistics Addition add new evidence-bearing material. That split turns out to be the story of the results.

What worked: evidence beats style, and both beat keywords

The three evidence-adding methods topped both metrics. Cite Sources, Quotation Addition and Statistics Addition achieved relative gains of 30 to 40 percent on Position-Adjusted Word Count and 15 to 30 percent on Subjective Impression. In absolute terms, Quotation Addition posted the single best score in the main table, 27.2 overall against a 19.3 baseline, with Statistics Addition at 25.2 and Cite Sources at 24.6. The interpretation the authors offer is straightforward: these edits give the answer-writing model verifiable, quotable, credibility-enhancing material, and the model rewards pages that supply it.

The second tier surprised people who expected substance to be everything. Pure presentation work, Fluency Optimization and Easy-to-Understand, produced visibility gains of 15 to 30 percent without adding a single new fact. The paper reads this as evidence that generative engines weigh how information is presented, not only what it contains. A clearly written page is easier for a summarizing model to lift from, so it gets lifted from more.

One representative example from the paper's qualitative table makes the scale vivid: on a Swiss chocolate query, simply attributing an existing consumption statistic to its source produced a 132.4 percent relative improvement for that page in that answer. Single examples are anecdotes, not averages, but the direction matches the aggregate result.

What failed: the keyword-stuffing irony

Keyword Stuffing, the one tactic in the lineup imported directly from old-school SEO, landed below the do-nothing baseline on the objective metric: 17.7 against 19.3. On Perplexity.ai it performed about 10 percent worse than baseline. Read that again in the context of who cites this paper. The study most often invoked to sell GEO services found that the most SEO-flavoured tactic in its test set made visibility slightly worse. The authors draw the obvious conclusion, warning that "techniques effective in search engines may not translate to success in this new paradigm."

Unique Words also did essentially nothing, drifting around baseline. And Authoritative, the persuasive-tone rewrite, produced no significant aggregate improvement, which the paper flags as a mildly counterintuitive finding: instruction-tuned models turn out to be fairly resistant to confident phrasing on its own. Sounding sure of yourself does not get you cited. Being checkable does. There is an old-advertising echo in that result which we unpack in our companion piece on Claude Hopkins and AI search: the researchers of 1923 and 2023 both found that specifics outperform superlatives.

The democratization finding: rank five gained most

Buried in Table 2 is the result we think matters most for independent businesses. When the authors optimized all sources simultaneously and broke results out by the page's original Google rank, the gains concentrated at the bottom of the retrieved set. With Cite Sources applied, the site ranked fifth gained 115.1 percent in visibility while the top-ranked site's visibility fell 30.3 percent. Quotation Addition and Statistics Addition showed the same shape, with rank-five gains of 99.7 and 97.9 percent respectively.

The mechanics explain it. Traditional rankings lean on backlinks and domain authority, which large incumbents accumulate and small operators struggle to match. But once the retrieval step has fetched its five sources, the generative model is choosing between passages of text, and text quality is a lever any creator can pull. A page that scraped into the retrieved set at position five can win the answer if its passages carry the best evidence. The authors frame this as GEO's potential to level the field between small creators and large corporations. For the local businesses we work with, that is the practical headline of the whole paper: the answer layer redistributes opportunity that the ranking layer had locked up.

Per-domain results: no single tactic wins everywhere

The aggregate table hides real variance, and the paper is explicit about it, "underscoring the need for domain-specific optimization methods." Table 3 maps each method to the query categories where it performed best. Statistics Addition led in Law and Government, Debate and Opinion queries, where data-backed assertions carry arguments. Quotation Addition dominated People and Society, Explanation and History, domains where a direct quote adds authenticity. Cite Sources performed best on factual and statement-style queries, where a verification trail matters most. Authoritative, weak in aggregate, still helped in debate-style and historical questions where persuasive writing has a natural home.

The combination analysis adds one more layer. Testing pairs of the top four methods on a 200-example subset, the authors found Fluency Optimization plus Statistics Addition beat any single strategy by more than 5.5 percent, and Cite Sources, middling alone in that subset, became strong in combination, averaging 31.4 percent improvement when paired. The working translation for a practitioner: write clearly, then load the page with sourced numbers, quotes and citations, and expect the exact mix that wins to depend on what kind of questions your customers ask.

The real-world check: Perplexity.ai

A fair objection to everything above is that it happened inside a research harness. The authors anticipated it and re-ran key methods against Perplexity.ai, a commercially deployed engine with a large user base. The pattern held. Quotation Addition again led on Position-Adjusted Word Count with a 22 percent improvement over baseline, Statistics Addition and Cite Sources improved up to 37 and 9 percent on the two metrics, and Keyword Stuffing again underperformed the baseline. One engine is not all engines, and 2023 is not 2026, but the replication is what separates this study from a demo: the same edits moved the same metrics on a system the authors did not build.

Why any of this works: the RAG architecture underneath

The GEO paper measures the effect. A 2020 paper from Facebook AI Research explains the cause. Lewis et al.'s "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (NeurIPS 2020) formalized the architecture that generative engines descend from: pair a language model's parametric memory, the knowledge baked into its weights, with a non-parametric memory, an external index of documents fetched by a retriever at answer time. Their experiments showed the hybrid generates more factual and more specific text than a model relying on its weights alone, and, critically, that you can change what the system knows by swapping the document index, no retraining required. In their world-leaders test, the same model answered correctly from whichever snapshot of the index it was handed.

That swappable external memory is the entire reason optimization is possible. If AI answers came purely from model weights, your content could not influence them between training runs; nothing you published on Tuesday would change Wednesday's answer. Because the engines instead retrieve live documents and ground their answers on them, the documents are an input you control. Aggarwal et al.'s methods are, mechanically, ways of making your document the one the generator leans on hardest once it has been retrieved. The RAG paper also foreshadows why evidence-rich pages win: grounding exists to reduce hallucination and give users something to verify, so material that is already sourced, quantified and attributed is the easiest raw ingredient for a grounded answer. We cover the business implications of that architecture in our explainer on what RAG means for your website.

What the paper does not claim

An honest teardown lists the boundaries, and the authors drew most of them themselves in their limitations section. Worth stating plainly:

Boundaries the citation-slingers omit

  • No traffic claim. Both metrics measure prominence within generated answers. Clicks, visits and revenue were never measured.
  • No guarantee. Results are averages across 1,000 test queries with per-domain variance the paper itself documents. Your page, in your category, may respond differently.
  • A point in time. The experiments ran on a gpt-3.5-turbo testbed and 2023-era Perplexity. The authors note the methods may need to adapt as generative engines evolve, mirroring the history of SEO.
  • No search-ranking evidence. The paper did not evaluate how these edits affect Google rankings, though it argues the changes are unlikely to hurt them since they alter text, not domains or backlinks.
  • Five retrieved sources. The testbed retrieves the top five Google results, so everything measured happens after retrieval. If your page is not being retrieved at all, these tactics have nothing to act on yet.
  • Benchmark drift. Query distributions change over time, and the authors acknowledge GEO-bench will need continuous updates to stay representative.

None of this weakens the paper. It positions it correctly: strong, replicated, directional evidence about which content properties generative engines reward, produced under stated conditions, with stated shelf-life. Our own cross-engine testing, written up in the AI engine consensus gap study, keeps confirming the corollary that engines built on different retrieval stacks disagree with each other, which is exactly what a domain-dependent, engine-dependent result set predicts.

How we run these findings in client work

Formative Digital's methodology assigns this paper's top findings to a named workstream: Vector 5, Cite. In practice that means every substantive page we build for a client carries attributed statistics with dates, quotations from credible primary sources, and outbound citations to the standards bodies, government data and academic work that validate the page's claims, the three edits the paper measured at 30 to 40 percent. The presentation findings feed Vector 4, Embed, where the same material gets written in plain, extractable sentences a summarizing model can lift cleanly.

"The reason we made citation work its own vector instead of a writing tip is this paper's Table 2. The rank-five site gained 115 percent when it added citations while the rank-one site lost ground, and that is the exact position most of our clients start from. When we rebuilt Mattress Miracle's content with sourced statistics and attributed quotes on every substantive page, we were running the paper's three winning methods as standing procedure, and the AI engines started returning the brand for queries where big-box retailers held every top ranking. You cannot out-backlink a national chain. You can out-evidence one."

Matt Griffin, Founder, Formative Digital

The usual YMYL caveat applies to that observation: Mattress Miracle's results (SEMrush, April 2026) reflect one engagement, and outcomes depend on industry, competition and existing digital presence. The paper's per-domain tables argue for the same humility, which is why our engagements start with testing which sources the engines already cite in the client's specific category rather than assuming the aggregate result transfers. Details on the full workstream live on our GEO services page, and the wider study library is in the research hub.

Get your free AI visibility audit

We will test where your pages currently stand in ChatGPT, Perplexity, Gemini and Google AI Overviews, and show you which of the paper's winning methods your site is missing. No charge, no obligation, a reply within one business day.

Read it yourself

The best outcome of this teardown would be more people reading the primary source. The paper is twelve pages, free on arXiv, and clearer than most of the commentary written about it. If you take one reading discipline from us, take this: when a vendor quotes the 40 percent figure, ask them which metric it refers to and which methods produced it. The ones who have read the paper will answer in a sentence. The ones who have not will change the subject, and that tells you what their research process looks like on your account, too. Questions about anything in this analysis, or about applying it to your own pages, are welcome through our contact page.

Frequently Asked Questions

What does the 40% visibility increase in the GEO paper actually mean?

It means the best-performing content edits raised a source's share of the generated answer, measured by Position-Adjusted Word Count, by up to about 40 percent relative to the unoptimized baseline. Position-Adjusted Word Count tracks how many words of the answer cite your page, weighted toward earlier placement. It is a share-of-answer metric computed inside a research testbed, not a measurement of clicks, traffic, revenue or Google rankings.

Which GEO methods worked best in Aggarwal et al.'s experiments?

Three content additions led the table: Quotation Addition, Statistics Addition and Cite Sources, each delivering roughly 30 to 40 percent relative improvement on Position-Adjusted Word Count and 15 to 30 percent on Subjective Impression. Two presentation edits, Fluency Optimization and Easy-to-Understand, also produced meaningful gains of 15 to 30 percent. Keyword Stuffing scored below the unoptimized baseline in the main experiment and about 10 percent worse than baseline on Perplexity.ai.

Does the GEO paper prove these tactics work on Google AI Overviews or ChatGPT today?

No. The paper tested a research generative engine built on gpt-3.5-turbo plus one deployed engine, Perplexity.ai, at a single point in time in 2023. The authors state plainly that methods may need to adapt as generative engines evolve, the same way SEO tactics evolved. The findings are strong directional evidence that citations, quotations and statistics raise answer visibility, but they are not a guarantee for any specific engine in any specific year.

Why do citations and statistics increase visibility in AI answers at all?

Because generative engines are retrieval-augmented systems. Lewis et al.'s 2020 RAG paper established the architecture: a retriever fetches a small set of documents and a language model writes the answer grounded on them. Once your page is inside that retrieved set, the model chooses which passages to lean on, and passages carrying verifiable evidence such as sourced statistics and attributed quotations give it more usable material. GEO tactics work on the selection step, not on the model itself.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '24), Barcelona. arXiv:2311.09735. All experimental figures in this article (GEO-bench composition, Table 1 to Table 5 results, limitations) are drawn from this paper. Link
  2. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020, Facebook AI Research. arXiv:2005.11401. Source for the RAG architecture, factuality findings and index hot-swapping result. Link
  3. Kumar, A., & Lakkaraju, H. (2024). Manipulating Large Language Models to Increase Product Visibility. arXiv:2404.07981. Cited by Aggarwal et al. as the adversarial counterpart to their non-adversarial optimization methods. Link
  4. ACM Digital Library. (2024). GEO: Generative Engine Optimization, KDD '24 conference record. doi:10.1145/3637528.3671900. Link