The 40% GEO uplift did not replicate
The paper everyone cites has been heavily qualified. Here is what survived, and what an agency should stop selling.
Almost every GEO pitch deck in circulation traces back to one 2024 paper reporting up to a 40% visibility uplift from adding citations, quotations and statistics to content. It is a good paper. It has also been substantially qualified by everything that came after it, and the qualification has not reached the decks.
What the follow-up work found
A 2025 benchmark tested 54 combinations of GEO method and domain. Three were significantly positive. More awkwardly, the gains it did find were largely zero-sum: when everyone applies the technique, the advantage disappears, because the technique was never creating value — it was redistributing attention among documents.
No content-side technique has yet demonstrated a stable causal effect on organic discoverability that holds under broad adoption. That is a genuinely hard thing to sell, and pretending otherwise is how this discipline earns the reputation SEO spent fifteen years escaping.
What did survive, with strong evidence
- Retrieval position and relevance dominate. Whether a document gets cited is mostly about whether it gets retrieved. Everything downstream is second-order.
- Off-site entity strength beats backlinks by roughly 3×. Brand mentions across the web correlate with AI visibility at about 0.66; backlinks at about 0.22.
- Around 84% of AI citations are earned or third-party media. Directories, review sites, comparison content, trade press.
- Freshness earns roughly 3.2× the citations. Recently updated content is retrieved disproportionately.
- No major AI crawler executes JavaScript. Content that only exists after hydration does not exist.
- `llms.txt` is effectively unused. Hundreds of requests across half a billion bot events, and Google states it ignores the file.
What that implies for an agency's roadmap
The high-confidence work is unglamorous and mostly not on the client's own website: get crawlable, get server-rendered, get into the third-party sources the engines actually read, keep things current, and build entity presence. The low-confidence work is the content tinkering that fits neatly into a monthly retainer.
This is exactly why we do not sell the intervention. A tool that scores you and sells you the fix has an incentive to keep the low-confidence work on the roadmap.
And measure it properly
If the effects are this contested, the measurement has to be good enough to detect one. A before-and-after on single-number scores from three runs cannot distinguish a real intervention from a Tuesday. Bands, run counts and fortnightly windows are not statistical pedantry here — they are the only way to know whether anything you did worked.
See it against your own client list
A working demo runs your prompts, in your market, on live engines — not a sandbox with seeded data. Bring one client brand and three competitors.