Research · Dispatch #9 ·
A citation is not proof: three ways AI search cites you and still gets it wrong
The AI-search industry treats one question as the whole game: was I cited or not? But a citation is a chain of promises: that the source supports the claim, that the address resolves to you, and that the version cited is your current one. Each link can break on its own, and across our own frozen-panel probing and a read of the field's foundational papers we have documented all three: the field's most-repeated numbers are not supported by the papers they cite; an engine rendered our exact content under a domain we do not own; and an engine cited our live URL while reproducing figures we had already corrected. None of these is common. The point is that 'cited: yes' is not a success metric, and only continuous first-party probing makes the failures visible.
The ninth GEO Glossary dispatch. (The sixth failed its evidence gate and was retired unpublished; we number by our project ledger, so the gap is deliberate, and one of the three findings below is the thing that dispatch tried and failed to claim.)
The AI-search visibility industry runs on one number: were you cited, yes or no. Tools count it, dashboards chart it, and a citation is treated as the win. This dispatch is about what that number hides. A citation is not a single event; it is a chain of promises. When an AI engine cites your page, it implies three separate things: that the source it points to actually supports the claim, that the address it emits resolves to you, and that the version it reproduces is your current one. Each of those links can break on its own, and a broken link still counts as "cited."
We have now documented all three, each from a different angle: one from reading the papers the field quotes, two from our own weekly frozen-panel probing. None of them is common. That is the honest headline, and it is the point: these are not epidemics, they are structural failure modes that the cited-or-not binary cannot represent, and each one is only visible if you look past the yes/no.
Link 1: the source does not support the claim
The first promise a citation makes is that the cited source says the thing it is cited for. At the level of the field's foundational literature, this one breaks routinely, and we checked it directly. The single most-repeated figure in GEO, "GEO can boost visibility by up to 40%," is a real sentence from a real abstract, but almost nothing else makes the trip into the marketing with it. The 40% is a position-adjusted word-share proxy measured on a 2023 GPT-3.5 testbed, and the same paper's real-search-engine headline quietly switches to a different, smaller metric. A separate multi-actor re-test found most content tactics ineffective or negative once everyone adopts them. And a verifiability audit of generative search engines found that only about half of AI-generated sentences are fully supported by their own attached citations.
That last number is the general case of this failure: roughly half the time, an AI answer's own citation does not fully support the sentence it is attached to. So the most basic promise, the source backs the claim, is the one with the most documented slippage, and it is slippage the reader cannot see without opening the source and reading it against the claim. "Cited" tells you a link exists; it does not tell you the link holds.
Link 2: the address is not yours
The second promise is that the emitted address resolves to you. This is the one we caught in our own data, and it is also the one we got wrong first, which is why it is worth telling carefully.
On one round in July, Microsoft Copilot rendered five of its citations to our brand and our exact page paths under domains we do not own: four under a hyphenated address that does not resolve at all, and one under a parked lookalike domain. The content was ours, the brand was ours, the path was ours; the address was not. A citation like that is "cited" by every counter, and it delivers zero traffic, because the URL goes nowhere.
Here is the honest part. When we first saw it, we drafted a dispatch about it. Then we ran a held-out round to check whether it persisted, and it did not: on the re-probe, none of the five rendered the phantom domain again, and the next several rounds recorded zero phantom addresses. So we retired that dispatch unpublished rather than generalize a single round into "engines fabricate domains." It was not a rate; it was an instability. Six weeks later, on one term in a late-August round, a wrong-domain render appeared once more, which is why this failure mode earns a mention here but not a claim of frequency. The honest description is bounded: engines occasionally emit a wrong or unresolvable address for content that is genuinely yours, we have seen it twice, and it is intermittent, not systematic. We log the rendered domain of every citation precisely so this stays an observation and never becomes an inference.
Link 3: the version is stale
The third promise is that the version cited is your current one, and this is the failure mode we found most recently and named: cited-version lag. An engine cites your page's live, current URL but reproduces a claim from an earlier version of that same page, one you have already corrected. The address resolves, the source is right, the content is genuinely something that page used to say. It just is not what the page says now.
We caught it on one of our own entries. We had corrected a set of figures on a page, and across two weekly rounds an engine cited that page's current URL while reproducing the pre-correction figures. We could rule out a simple misreading of the live page, because the specific numbers the engine returned appear nowhere on the current version; they only existed before the fix. The behavior is consistent with a stored, stale copy of the page, which is a documented retrieval reality: an independent investigation of one major engine's retrieval stack found a shared reading cache that serves stored page copies, some observed served more than ninety days after they were fetched, and that engine's own vendor documentation notes that indexed or cached pages can be older than the live version.
The uncomfortable property of this one is that it is nearly invisible. You can only detect it where the engine reproduces the exact fact you changed, and most of the time engines reproduce a page's framing, not its specific numbers, and the framing does not change when you fix a figure. So our evidence is a single term on a single engine across two rounds: enough to name the failure mode, not enough to say how often it happens. We report it as a bounded observation, not a rate, and the honest consequence is that the cases we cannot see are unknown, not assumed.
What this means if you measure AI citation
- Treat "cited" as the start of the check, not the end. A yes on the inclusion question leaves three things unverified: does the cited source support the claim, does the address resolve to you, and is the version current. Each is a separate look, and each can fail while the citation counter reads success.
- Read the rendered URL, not just the brand. A citation that names you under a domain you do not own delivers nothing; a visibility tool that credits brand mentions will over-count it. Check that the address resolves to your site.
- Treat a correction as landed only after you verify it. After you fix a specific fact, re-query that fact and read whether the old or new value comes back. A clean-looking answer that reproduces only your framing is not evidence the correction propagated.
None of this requires you to distrust AI citations wholesale; most citations are fine. It requires you to stop treating a citation as self-certifying. The three failures share one thing: they are only visible from continuous, first-party probing of the actual answer interface, checking the source, the address, and the version of each citation over time. A vendor selling a "were you cited" dashboard cannot see them, because the dashboard is built on the exact binary that hides them.
Limits, honestly
Two of these three failures rest on thin evidence, and we would rather say so than round up. Link 2 (the wrong address) is two observations, intermittent, and we killed our own earlier dispatch for overclaiming it. Link 3 (the stale version) is one term on one engine across two rounds, and it is structurally hard to observe, so we make no claim about how common it is. Only Link 1 (the unsupported claim) rests on published, at-scale audits, and even there the general "about half" figure is one study's measurement, not a law. What ties the three together is not a shared frequency but a shared blind spot: each is a way a citation that "happened" is nonetheless wrong, on an axis the yes/no metric does not have. The value of naming them is that a name is what lets you go looking, and a field that measures only inclusion has not been looking.
More dispatches
- Dispatch #8We built this glossary on coining our own terms. Does AI search actually reward it?
- Dispatch #7We refuse to invent benchmarks. Does AI search punish us for it?
- Dispatch #5GEO's most-cited numbers, checked against the papers they come from
- Dispatch #4One panel, five engines, mostly separate citation sets: 'cited by AI' is not one thing
- Dispatch #3Cited more on Gemini, less on ChatGPT: a gradient, and what it is not
- Dispatch #2Google caught up: the AI-citation gap looks like a reporting lag
- Dispatch #1AI engines cited this page before Google indexed it