Research · Dispatch #7 ·

We refuse to invent benchmarks. Does AI search punish us for it?

Several of our glossary entries decline to give the number people search for: no target citation match rate, no standard attribution rate, no proven lift from authoritative tone. We tested whether that honesty costs us citations by asking five AI engines the number-seeking version of each question and comparing against the definitional control. Across 16 paired probes where our page held the definitional citation, it kept the citation 11 times, and not one loss went cleanly to a page that asserts a number. On the engines that carry editorial framing, the answers adopted our hedges and debunked the inflated figures. Honesty held where we hold the term.

The seventh GEO Glossary dispatch. (The sixth, a Copilot domain-misattribution finding, failed its held-out persistence gate and was retired unpublished. We number dispatches by our project ledger, not by what survives review; a numbering gap is what an evidence gate looks like from outside.)

This one is about a policy of ours that looked like a liability. Several of this glossary's metric entries refuse to give the number people actually search for. Citation match rate says the metric has no universally accepted target and that any threshold is a practitioner heuristic. Attribution rate says no vendor or academic literature defines a canonical formula, and reports per-engine ranges instead of a goal. Authoritative statement strength reports the honest, deflating number: in the original GEO benchmark, an authoritative-tone rewrite produced no statistically significant improvement, and the paper says so verbatim. Meanwhile, plenty of pages in this space will happily tell you to aim for a 25 percent lift, a 30 percent attribution target, a stage-by-stage benchmark table.

So the worry writes itself. When a user asks an AI engine "what is a good citation match rate to aim for," the engine wants to hand back a number. We decline to assert one. Does the citation go to whoever asserts one?

We had a concrete reason to take the worry seriously. On a sibling site we also run (an exam-prep tool in a crowded education niche), the page that honestly says "there is no official score threshold" lost exactly this way: it had been cited by three engines on the definitional question, then dropped to zero on the number-seeking question, which went to competitors who assert a pass mark. Honesty, on that site, was intent-conditional: it helped on definitional queries and lost the give-me-a-number ones.

The experiment

We ran the direct test on this glossary, twice. The design is a paired probe. For each hedged term, we ask the frozen definitional question our weekly panel always asks ("What is citation match rate in AI search?") as the control, and a number-seeking variant ("What is a good citation match rate to aim for in AI search?") as the treatment, in the same round, on the same five engines (ChatGPT, Perplexity, Claude, Copilot, Gemini), with the same source-request suffix. The interesting cell is the one the worry predicts: the definitional prompt cites us, the number-seeking prompt does not, and the winner is a page that asserts a number.

Round one, in late July, paired two terms. Round two, on August 10, widened it to five hedged metric terms (citation match rate, authoritative statement strength, attribution rate, citation share, citation velocity), each verified beforehand to genuinely decline a target number on the page. Five terms times five engines is 25 pairs; in 16 of them the definitional control cited us, so 16 pairs could test the hypothesis at all.

What happened

Of the 16 pairs where we held the definitional citation, the number-seeking variant kept citing us in 11. And in zero of the five losses did the citation go cleanly to a number-asserting page.

The holds concentrate exactly where our authority does. Claude held all five of its pairs. Perplexity held all three of its pairs, and added a bonus in the wrong direction for the hypothesis: on citation share, where the definitional prompt did not cite us, the number-seeking prompt did. ChatGPT held three of four; its one loss (attribution rate) went to an independent measurement-research page, not to anyone asserting a benchmark.

The five losses live on Gemini and Copilot, and they dissolve on inspection. Gemini's two look like the hypothesis until you check its definitional controls: one of the two, probed three times in the same session, went cited, not-cited, not-cited. An engine that flips two out of three times on the identical frozen prompt has an inclusion-volatility floor higher than any honesty effect we could measure through it. Copilot's two losses have a different texture: its definitional answers for both terms were built almost entirely on our pages (one was literally a single-reference answer, ours), and the variant then rotated to a how-to blog. That is retrieval churn on a references-list engine, not a penalty for hedging. The single cell all round where a number-asserting page did win a variant we lost sits on that flipping Gemini control, so we do not count it as the hypothesis firing.

What the winning answers did with our content is the part we did not predict. On the engines that held, the number-seeking answers did not treat the hedge as a gap to route around. They adopted it as the frame. Perplexity's answer to "what is a good citation match rate" opens by stating there is no universal good rate yet, reproduces our warning that the denominator must be fixed before comparing tools, offers its own 80 and 90 percent working bands, and then explicitly labels them operational targets rather than industry standards. Its answer on authoritative tone goes further and debunks the going figures: "Do not plan on a fixed lift such as 25%, 30%, or 40%," on the grounds that those numbers come from marketing articles and are not supported by the primary benchmark, citing the original paper and our entry. ChatGPT's version opens with "Very little, based on the best available evidence" and prints the honest results table. Claude's answers repeatedly quote a sentence we added to the citation-match-rate entry one week earlier, cautioning against chasing one universal target. The number-seeking queries, on these engines, produced sourced debunks of the very numbers we refused to invent.

The reconciliation with the site that lost

Two sites, the same honest-hedge editorial policy, opposite outcomes on number-seeking queries. The variable that reconciles them, on our current evidence, is authority over the term. On this glossary, the tested terms are practitioner-coined measurement vocabulary where our entries are the closest thing the query has to a primary source; an engine that wants to answer carefully has to reckon with our framing. On the exam-prep site, the honest page is one voice in a crowded niche with a genuine canonical owner (the education department) and many competitors willing to assert a pass mark; there, the honest page is optional, and on the number-seeking query it lost.

So the working rule we take from this is conditional, not triumphant: an honest hedge survives number-seeking queries where you are the authority on the term, and is at risk where you are one contested voice among many. Honesty is not an independent citation moat. It is a benefit you can bank once the territory is yours, and a real risk where it is not.

Two refinements ride along. First, the effect is engine-architectural. The engines that held (Claude, Perplexity, ChatGPT) are the ones that transmit editorial framing into their answers; the engines where pairs failed (Copilot, Gemini) select references in ways that look more like retrieval rotation, and their definitional citations were unstable to begin with. Second, a number-seeking query can actively help an evidence-dense honest page: on authoritative statement strength, the plain definitional prompt lost to generic official documentation on two engines in round one, while the number-seeking variant surfaced our page, because ours is the one carrying the actual measured number, deflating as it is.

What this means if you run a content site

If your pages hedge honestly on questions where users want a number, the practical questions are where you can afford it and what to pair it with. On terms where you are the primary or originating source, our evidence says hold the line: the engines that carry framing will use your caveats as the skeleton of their answer, and the inflated numbers get debunked with your page as the source. On contested terms with a stronger canonical owner, an honest hedge alone is exposed; whatever real evidence you do hold (a measured range, a dated study, a named condition) is the part worth leading with, because specific empirical content is what pulled our pages into number-seeking answers. And read results per engine: a blended visibility score would have averaged Claude's five holds against Gemini's coin-flips and told you nothing.

This dispatch is also the other half of an earlier one. Dispatch #5 audited how the field's foundational numbers get quoted with their conditions stripped. The supply side of that problem is pages asserting figures the underlying research does not support; this dispatch is what the demand side looks like at the answer layer, and, at least where we hold the term, the engines are on the side of the conditions.

Limits, honestly

Twenty decisive pairs across the two rounds (sixteen of them in the wider second round) is a small sample, and the variant prompts are one wording each, not themselves frozen-tested. The engine-architecture split (framing-transmitting versus references-list) is our post-hoc reading of five losses, not a pre-registered hypothesis. The two-site contrast is suggestive, not causal: the sites differ in niche, competition, and age, not only in term authority, and we run both, so this is a replication across our own properties, not an independent one. Claude changed underlying models between the two rounds, which flatters its round-two breadth. And a citation held this week is not a citation held forever; the same instrument that produced this result exists because these things churn. We will keep running the pairs.

More dispatches