GEO: Generative Engine Optimization
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande. Submitted November 2023, accepted to KDD 2024. The paper that named this wiki’s central concept, and until now the wiki did not hold it — generative-engine-optimization was assembled entirely from trade press written about it.
What it contributes that the trade press does not
A metric — though see the correction below before trusting it. The spoke’s standing complaint is that GEO advice is unfalsifiable because nobody measures recommendation share; every number in this corpus has been a click number, which is the wrong instrument for a channel where there is often no click. Aggarwal et al. define visibility over the generated answer itself — a position-adjusted word-count measure of how much of the response a source occupies, plus a subjective-impression measure — and optimize against it. That is a consideration metric, and it is the first one here.
A benchmark. GEO-bench: roughly 10,000 queries across nine datasets and multiple domains, which makes the claims reproducible in a way no agency case study is.
Results with a shape. Visibility gains up to 40%, with the largest coming from adding verifiable statistics, adding credible quotations, and citing sources — around 30–40% on the position-adjusted metric. Nine optimization methods were tested and the paper reports that efficacy varies by domain, which is a finding the “best practices” genre routinely flattens.
Where it lands against what this wiki already believed
brand-depth-ai-recommendations argues that entity salience and relationship density drive recommendations and that citations are outcomes rather than levers. The GEO paper cuts across that: the interventions it measures are content-level and local — put a statistic in this passage, add a quotation to that one — and they moved a metric. Both can be true (a brand-level prior and a passage-level edit are different mechanisms) but the corpus should stop treating the brand-depth account as the whole story, because the only controlled measurement here is of the other kind.
It also gives the ai-search-shift thread its missing baseline. Trade sources describe a channel where measurement is impossible. This paper measured it in 2023 and published the benchmark.
Limits, and one to chase
The visibility metric is a proxy: occupying more of an answer is not the same as being recommended, and neither is the same as a sale. The paper optimizes against generative engines the authors constructed for the benchmark, not against ChatGPT’s or Google’s production systems, so transfer to the live platforms is assumed rather than shown — and those platforms change under you.
Read properly the same day, and it does not survive intact. c-seo-bench (Puerto et al., NeurIPS 2025) evaluated these methods across two tasks, six domains and varying numbers of competing actors, and found them “largely ineffective” and often harmful — the Statistics method, one of the best performers above, decreases rankings in 19 of 24 settings. The authors trace the disagreement to the metric: word count “does not measure the LLM preference, contrary to the citation ranking that we use.”
So read this paper’s 40% as share of answer text, which is what it measured and what it says it measured, and not as recommendation or citation preference. Its lasting contributions are the framing, the vocabulary and GEO-bench; its headline effect did not replicate.
Related
generative-engine-optimization · ai-search-shift · brand-depth-ai-recommendations · c-seo-bench · ai-content-seo-visibility · answer-engine-optimization · synthesis