A critical survey of generative engine optimization, posted to arXiv on July 15, 2026, concludes that no reviewed technique produces a stable, cross-platform effect on organic discoverability or downstream traffic, directly undercutting the "40 percent" visibility figure that a fast-growing optimization industry has built its pitch around.
The paper, titled Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization, was written by Olivier Martinez and submitted to the Information Retrieval section of arXiv under identifier 2607.14035. It reviews 45 studies published between November 16, 2023, the date of the first version of the foundational GEO paper, and July 14, 2026, spanning peer-reviewed articles, papers accepted but not yet presented, workshop contributions, and preprints, and it grades each by evidentiary weight rather than treating them as interchangeable.
Its central argument is narrow and specific. According to the survey, experiments provide strong evidence that a document already placed in an engine's context can change how that document is cited or used, but they establish far less often that a page will be retrieved at all, and almost never that any of this produces a durable effect on clicks or conversions. That is the difference between a laboratory result and a business outcome, and the paper argues the two have been routinely conflated.
The 40 percent figure the survey rejects
The most widely circulated number in generative engine optimization comes from the foundational paper by Aggarwal and colleagues, published at the ACM SIGKDD conference in 2024, which reported gains of up to roughly 40 percent. The survey traces the figure to a single metric in a single configuration.
According to Martinez, the 40 percent derives from an increase in Position-Adjusted Word Count, a measure that discounts passages appearing later in an answer, from 19.3 to 27.2 under a strategy called Quotation Addition, which works out to approximately 41 percent in relative terms. That gain occurs inside a testbed where five documents have already been supplied to the generator. It does not mean 40 percent more readers will click, nor that a page gains 40 percent in the probability of being retrieved. It means that, in that specific setup, a source already handed to the model receives a larger position-weighted share of the attributed text.
The survey places the claim "GEO increases visibility by 40 percent" in its lowest confidence tier, rejected as a general statement and described in a summary table as a relative maximum on one metric under a specific configuration.
That correction matters because the number has traveled far from its source. A Wall Street Journal investigation, covered by PPC Land, documented that businesses now pay substantial sums to shape how ChatGPT and other assistants describe them, planting brand authority statements across multiple sites and using superlatives to trigger the systems. In that reporting, AI chatbot referrals had grown from under a million visits in early 2024 to more than 230 million monthly by September 2025, and accounted for 44 percent of traffic at some optimization firms.
Visibility is a vector, not a rank
The paper's framing device is a rejection of the single-number scoreboard. Rather than treating AI visibility as one rank to be won, Martinez proposes a visibility vector that separates several distinct quantities: discoverability, or the probability a page is retrieved at all; exposure within the context window; the probability of a citation; observable prominence; absorption, meaning the source's actual contribution to the facts and wording of an answer; fidelity, the extent to which attributed claims are genuinely supported; and the behavioral or economic outcome, such as a click or a conversion. The point of the separation is that an intervention can improve one component while damaging another. Collapsing these into one score, the survey argues, hides the trade-off and lets a promotional metric stand in for a business result.
This maps onto instability the marketing industry has already measured. Analysis covered by PPC Land found that Google exchanges 56 percent of its AI Mode citation sources weekly while ChatGPT exchanges 74 percent, churn that makes any single week's citation data a shaky basis for strategy. The survey formalizes why: a point estimate taken on one day, for one query phrasing, on one engine, is not a stable indicator.
When the rewrite backfires
The most operationally pointed result concerns what happens when optimization is tested across the full pipeline rather than inside a fixed context. According to the paper, an end-to-end environment called SAGEO Arena, run across 171,003 documents and 2,700 queries, found that optimizing only the body of a page reduced average top-20 presence by roughly 9 percent, cut top-10 presence after reranking by 16 percent, and lowered final citation by 6 percent. Applying one automated optimization system to the body alone produced larger losses still.
The mechanism is simple once the stages are separated. If a transformation raises the conditional probability of citation given retrieval but lowers the probability of retrieval, the total effect can be negative. A page can be rewritten to perform better once selected while becoming less likely to be selected at all, a danger invisible to any test that starts with the document already in context.
Generic heuristics travel poorly
A recurring finding across the reviewed corpus is that the transferable "tips" central to much commercial GEO advice do not generalize. According to the survey, a benchmark called C-SEO Bench tested optimization methods across two tasks, six domains, roughly 1,900 queries, and 16,360 documents, and found only three of 54 method-domain combinations significantly positive in the main experiment, with none positive in question answering. Several transformations reduced rank outright. The paper reports a parallel result in e-commerce, where ten of fifteen initial heuristics were neutral or negative and only systematically optimized prompts performed better. It also notes that keyword stuffing, imported from conventional search, does not transfer and can lower the position-adjusted metric.
Two factors do hold up. According to Martinez, query-document relevance is the most reproducible lever, and position within the context is close behind. A large factorial experiment cited in the survey, comprising 252,000 trials across six language models and eighteen factors, identified relevance and position as the primary determinants of the first citation. Extractable evidence such as verifiable figures, definitions, dated facts, and prices shows a more moderate advantage, and the effect depends on intent: a recent date helps a time-sensitive query but not a stable definition. The survey notes that a fabricated statistic may increase reuse while degrading the accuracy of the answer, so the operative criterion is verifiable, dated, properly attributed evidence rather than simply "add numbers."
The industry has been moving toward the same conclusion from the practitioner side. A scored analysis of 54 studies by Cyrus Shepard, covered by PPC Land, placed URL accessibility and search rank as the two highest predictors of AI citation, at 9.5 and 9.4 on an evidence-weighted scale, pointing back toward traditional retrieval fundamentals rather than novel formatting tricks. Google's own position, published in official documentation and reported by PPC Land, is that optimizing for generative AI search is optimizing for the search experience and therefore remains search optimization.
Engines do not share a source list
One of the survey's firmer conclusions is that there is no single global ranking to optimize for, because different generative surfaces cite different sources. According to the paper, an audit of 4,706 queries found that 53 percent of domains cited by Google's AI Overviews did not appear in the organic top 10, and 27 percent were absent from the top 100. A separate audit across 11,500 queries reported URL-level overlap between organic Google, AI Overviews, and Gemini in the range of 0.11 to 0.18.
Activation rates compound the fragmentation. The survey cites a study of 55,393 trending queries over 40 days in which the overall AI Overviews activation rate was 13.7 percent but rose to 64.7 percent for queries phrased as questions, while a different representative sample showed AI Overviews on 51.5 percent of queries. The figures are not contradictory; they reflect different query distributions, which is why a rate quoted without a description of its sample carries little weight.
This cross-engine divergence is well documented in trade coverage. Research covered by PPC Land, examining 89,000 LinkedIn URLs cited across ChatGPT Search, Google AI Mode, and Perplexity, found that Perplexity cites company pages 59 percent of the time while ChatGPT Search and Google AI Mode cite individual creators 59 percent of the time, so a brand publishing only on its company page would reach one surface and be largely absent from the others.
Citation does not mean support
The survey draws a sharp line between being cited and being accurately represented. According to a foundational fidelity study it reviews, across four early generative engines only 51.5 percent of sentences were fully supported by their cited sources, and 74.5 percent of citations actually supported the proposition they were attached to. Newer work reports credible-source shares from 71.4 to 86.3 percent depending on the assistant and topic, and classifies roughly 11 percent of 98,020 atomic claims as insufficiently supported. There is also the question of what is being cited: one audit found that approximately 16 percent of more than 19,000 retrieved textual pages were labeled AI-generated by the detector used, while 27.1 percent of URLs could not be scraped at all because they were inaccessible, removed, or non-textual. A visible citation, the paper concludes, may point to a page that does not support the claim, or to synthetic content feeding back into the system.
Traffic evidence is the weakest link
For all the money flowing into generative engine optimization, the survey finds the evidence on actual traffic and revenue to be the thinnest in the field. According to Martinez, one log-based natural experiment found that total ChatGPT referrals to a site increased by a factor of 5.7, but untreated pages on the same site had already risen by a factor of 3.5 as the platform itself grew. After controlling for that growth, the estimated additional effect was a multiplier of 1.82, with a 95 percent confidence interval from 1.31 to 2.54, and a conservative placebo test yielded a statistically inconclusive result. A separate industry study reporting a 20 percent production traffic lift is noted as rare real-world evidence, but the survey observes that its group sizes and intervention components were not described in enough detail to support a general estimate. Its blunt summary is that claims about return on investment from GEO clearly outstrip the academic evidence.
That gap is consequential given how much traffic is genuinely at stake in the shift to AI answers. A randomized field experiment involving 1,065 desktop Chrome users, covered by PPC Land, produced the first causal evidence that Google's AI Overviews cut outbound organic clicks by 39.8 percent and raised zero-click searches by 34.5 percent, with no measurable improvement in how users rated their search experience. Separate research tracked by PPC Land found that when AI Overviews appear, click-through rates at position one fall from 27 percent to 11 percent, and a February 2026 Ahrefs study of 300,000 keywords found AI Overviews correlated with a 58 percent reduction in click-through rates for top-ranking pages. The traffic loss is real and measured; the survey's point is that the case for GEO as the remedy is not.
Optimization and manipulation share a channel
A structural theme of the paper is that white-hat optimization and outright manipulation operate through the same mechanism, since both modify the text an engine consumes. According to the survey, the normative distinction should rest on truthfulness, semantic preservation, disclosure, and the absence of hidden instructions rather than on visibility alone. It proposes four cumulative tests for a legitimate rewrite: that facts remain true, that statistics and references are verifiable, that the document informs the user rather than issuing hidden commands to the model, and that commercial intent is disclosed without fabricated disparagement of competitors. The risk is not hypothetical. The survey documents that indirect injection into a document can raise a target by roughly three ranks on one commercial search API, and that preference-manipulation attacks increased the recommendation rate of a fictitious camera from 34.0 to 59.4 percent in one study, with some plugin selections rising by as much as a factor of 7.2.
This connects to a live tension over whether aggressive tactics are already being treated as spam. Coverage by PPC Land of Lily Ray, of the agency Amsive, documented her warning that penalties now compound across traditional and AI search surfaces at once, so a site penalized for aggressive optimization is also removed from the pool AI systems draw on when generating answers. The survey argues that defenses which are too permissive reward high-volume content producers while filters that are too strict penalize small publishers describing their own products.
Why the survey matters for marketers
The document arrives as marketing budgets are being redirected toward generative visibility on the strength of numbers the survey argues cannot bear the weight. IAB Slovenia, in research covered by PPC Land, now lists GEO alongside conventional search optimization as a distinct tracked activity cluster, a sign of how quickly it has moved from concept to budget line. Adobe research reported by PPC Land found that 98 percent of marketers lack a clear roadmap and full confidence in their AI optimization approach.
The survey's contribution is not a set of tactics but a standard of proof. It proposes that a study state up front whether it is measuring the effect of a rewrite once a document is injected, the total effect in a reproducible pipeline, an observational association with a commercial surface, or a business outcome in production, because conflating these yields invalid conclusions. It recommends testing across multiple named engines with version and date recorded, multiple query phrasings, and an untreated baseline, and it notes that outputs containing no search or no citations are data, not results to be discarded.
The measurement discipline is grounded in observed instability. According to the survey, one audit found that repeated runs at a temperature setting of zero, which is supposed to make output deterministic, still changed between 9 and 28 percent of decisions on the surfaces where that setting could be controlled, and in one configuration studied, 57.8 percent of ChatGPT repetitions did not activate web search at all. A dashboard that computes a share of citations only among responses that contain citations, then presents that as overall visibility, is measuring the wrong denominator.
The paper's closing position is that generative engine optimization is a real phenomenon with a real, if narrow, established effect, and that the work remaining is to turn it into a cumulative discipline rather than a promotional vocabulary. Within the reviewed corpus, the best-supported synthesis is that already-retrieved content can causally influence an answer, while no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream clicks and conversions.
Timeline
- November 16, 2023: First arXiv version of the foundational GEO paper by Aggarwal and colleagues is released, opening the survey's review window.
- 2024: The foundational GEO paper is published at the ACM SIGKDD conference, introducing the "up to 40 percent" figure and the GEO-bench benchmark of 10,000 queries.
- November 16, 2023 to July 14, 2026: Publication window covered by the survey, spanning 45 reviewed studies.
- July 14, 2026: End of the survey's review window.
- July 15, 2026: Olivier Martinez submits "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization" to arXiv as identifier 2607.14035, an 18-page paper with 8 tables and 1 figure.
- July 20, 2026: The SIGIR 2026 conference begins, the venue for two studies the survey lists as forthcoming.
Related PPC Land coverage
- How brands manipulate ChatGPT to dominate AI search results documents the Wall Street Journal investigation into paid tactics for shaping AI recommendations and the growth of chatbot referral traffic.
- AI Overviews cut publisher clicks 39.8% in first randomized study reports the first causal measurement of how AI Overviews divert clicks away from publisher sites.
- 23 factors that actually get your content cited by AI search engines covers Cyrus Shepard's scored analysis of 54 studies ranking the predictors of AI citation.
- Lily Ray: what the SEO industry is getting dangerously wrong about AI search examines how aggressive optimization tactics can trigger penalties that compound across traditional and AI surfaces.
- Google's VP of Search tells CMOs: good SEO is still all you need for AI reports Google's official position that optimizing for AI search remains conventional search optimization.
- Semrush maps how LinkedIn content earns citations in AI search tools details platform-level differences in which sources ChatGPT Search, Google AI Mode, and Perplexity cite.
- What Ahrefs' fake brand experiment actually proved about AI search covers a controlled test of how AI systems handle brand claims and manipulation.
- Original research tops AI search citations at 82%, NP Digital survey finds reports on content types most associated with AI citation and marketer readiness for the shift.
Summary
Who: Olivier Martinez, affiliated with Sciences Po, authored the critical survey; it reviews work from researchers across the generative engine optimization field, including the foundational paper by Pranjal Aggarwal and colleagues.
What: A critical survey of 45 studies concluding that no reviewed GEO technique produces a stable, cross-platform effect on organic discoverability or downstream traffic, and that the widely cited "40 percent" visibility gain describes a relative maximum on one metric in one fixed-context configuration rather than a general result.
When: The paper was submitted to arXiv on July 15, 2026, covering research published between November 16, 2023, and July 14, 2026.
Where: Posted to the Information Retrieval section of arXiv as identifier 2607.14035; the reviewed studies span venues including ACM SIGKDD, EMNLP, ICLR, ACL, and SIGIR, alongside numerous preprints.
Why: The survey argues that commercial claims about generative engine optimization have advanced faster than the evidence supporting them, and it offers a formal model, a visibility vector, an evidence hierarchy, and a measurement protocol intended to separate demonstrated effects from promotional figures as marketing budgets shift toward AI visibility.
Discussion