C-SEO Bench: Does Conversational SEO Work?
Haritz Puerto, Martin Gubri, Tommaso Green, Seong Joon Oh and Sangdoo Yun — Parameter Lab, UKP Lab (TU Darmstadt), University of Mannheim, University of Tübingen and NAVER AI Lab. NeurIPS 2025 (arXiv 2506.11097v3, 20 October 2025).
The answer in its title is no, and it is aimed directly at geo-paper.
What it tested
The first benchmark to evaluate C-SEO methods across multiple tasks, domains and numbers of competing actors: two search tasks (question answering, product recommendation) with three domains each, plus a protocol that varies how many content owners adopt the technique at once. Prior evaluations, including the GEO paper’s, tested a single actor in a narrow domain.
Three findings, each of which costs this wiki something
The methods mostly do not work, and often backfire. “Most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected.” Concretely: the Statistics method — one of the GEO paper’s best performers — decreases rankings in 19 of 24 evaluated settings. On product recommendation with Haiku 3.5, 26 of 30 cases show significant negative effects; on question answering with o4-mini, 19 of 30.
The reason is the metric, and this is the important part. The authors are explicit about why they disagree with Aggarwal et al.:
“The differences are due to the metrics used. They report their main results as word count, i.e., the ratio of words used in the LLM response to talk about the target document. However, this metric does not measure the LLM preference, contrary to the citation ranking that we use. A higher word count does not necessarily correspond to a better citation ranking.”
What actually moves visibility is old-fashioned retrieval ranking. “The initial ranking of documents retrieved by the search engine plays a far more dominant role.” Traditional SEO is significantly more effective than any C-SEO method tested, so C-SEO complements SEO rather than replacing it.
And it is zero-sum. As adoption rises, per-actor gains fall — “a congested and zero-sum nature of the problem” — the same dynamic classic SEO produced.
Correcting this wiki, same day
geo-paper was ingested here on 2026-08-08 as closing the spoke’s #1 growth edge: a consideration metric measuring recommendation share rather than clicks. That was wrong, and this source is the reason.
Position-adjusted word count is not a consideration metric. It measures how much of an answer a source occupies, and C-SEO Bench shows that quantity moves independently of whether the model actually prefers and cites the source. So the edge is not closed: the corpus acquired a metric and then learned from a peer-reviewed benchmark that the metric does not measure the thing.
Both papers stay. GEO is still the founding work and still the source of the technique vocabulary; this is the multi-domain, multi-actor replication that did not reproduce it and explained the discrepancy. That is how the disagreement should be read — not as one paper being discredited, but as a single-actor result in a narrow domain failing to generalise, measured on a proxy that turned out to be loose.
What it means for the practice this wiki documents
Most GEO advice in the corpus descends, directly or through trade press, from the interventions this benchmark just found ineffective or harmful. generative-engine-optimization should be read with that in front of it. The advice most likely to survive is the least novel: be retrievable, rank well, because retrieval ranking dominates everything downstream of it.
Related
geo-paper · generative-engine-optimization · ai-search-shift · brand-depth-ai-recommendations · answer-engine-optimization · synthesis