Research · Dispatch #11 ·

We corrected a number, and an AI engine served the old one for five weeks

On 2026-08-04 we corrected a set of figures on one of our pages. One AI engine kept citing that page's live URL while reproducing the pre-correction figures for five consecutive weekly probes, then rendered the corrected figures on the sixth, about 48 days later, with no action from us. This is cited-version lag measured end to end on a consumer engine: a versioned correction plus a weekly directed re-probe of the exact changed fact. It is a single case (one correction, one engine, one page), so the number is not a rate; but the shape (persist for weeks, then self-heal) and the method are the point. The practical rule that falls out: after you fix a specific fact, treat the correction as submitted, not landed, and re-probe rather than re-edit.

The eleventh GEO Glossary dispatch. (We number by our internal ledger, so there are gaps: #10 was scoped and then failed its evidence gate, so it was retired unpublished, the same reason #6 is missing. We publish the ones that survive their evidence check.)

On 2026-08-04 we corrected a set of figures on one of our own pages. For the next five weekly probes, one AI engine kept citing that page's current URL while reproducing the old figures, the ones we had already replaced. On the sixth probe, about 48 days after the fix, it finally rendered the corrected numbers, with no further action from us. This dispatch is that episode, measured end to end.

We named this failure mode in Dispatch #9 as one of three ways a citation can be "yes" and still wrong: the version axis, where the citation points at the right source at a resolving address, but the version the engine reproduces is stale. Dispatch #9 named it; this one measures a single instance of it from correction to clean render. The term is cited-version lag.

The one thing to hold before any number: this is N=1

One correction, one engine, one page. Everything below is a measured single case, not a rate and not a general refresh time. We are not claiming AI engines take 48 days to update; we are reporting that this one did, and showing the method that let us watch it. The shape generalizes; the number does not.

The method: a versioned correction plus a weekly directed re-probe

Cited-version lag is nearly invisible by default. You only catch it where an engine reproduces the exact fact you changed, and most of the time engines reproduce a page's qualitative framing (its argument, its definition) rather than its specific figures, and that framing does not change when you fix a number. So detecting it takes a deliberate setup:

  1. Make a correction that changes a specific, verbatim figure, with a dated timestamp on the page.
  2. Each week, re-query the exact fact on the engines that render specifics, not just the framing.
  3. Record whether the answer shows the old value or the new one.

Our correction was on the authoritative statement strength entry. The figures come from the GEO benchmark paper (Aggarwal et al. 2023); we had originally cited the paper's plain Word Count column (baseline 19.5, Authoritative 21.8, a +11.8% relative gain, Quotation +42.6%) and on 2026-08-04 corrected them to the paper's position-adjusted "Overall" column (baseline 19.3, Authoritative 21.3, +10%, Quotation about +41%). The load-bearing finding never changed (the paper calls Authoritative tone "no significant improvement" either way); only the raw figures moved. That gave us a clean before/after pair to watch, because the pre-correction relative-lift set appears nowhere on the current page. If an engine reproduced +11.8% and +42.6%, it was reading a stored copy, not misreading the live one.

The curve

We probe five engines weekly on a frozen panel. Only one, Claude, renders the specific PAWC figures; the others reproduced the qualitative "no significant improvement" framing and so were silent on the version question either way (more on that below). Here is what Claude showed, week by week:

date day rendered for the corrected fact
2026-08-04 0 figures corrected on the live page
2026-08-17 13 cites the current URL, reproduces the pre-correction figures (stale)
2026-08-24 20 stale, second round
2026-08-31 27 stale, third round
2026-09-07 34 stale, and a same-session triple re-probe returned stale, then current, then stale
2026-09-14 41 stale, fifth round
2026-09-21 48 resolved: renders the corrected figures

Two things a single sighting could never show:

It flickered before it healed. On day 34, the same prompt run three times in one session returned the stale render twice and the corrected render once. So the lag was not a fixed state that flips over cleanly; in that session the stored copy was sometimes served and sometimes bypassed. That is consistent with a cache or retrieval layer that is intermittently hit, not a switch.

It healed on its own. We never re-edited the page. Re-editing would have been the wrong move (the page was already correct), and it did not need it: at some point between day 41 and day 48 the engine's representation caught up.

Why it was even visible, and why 48 is a lower bound

Across the same six weeks, the four other engines that cited the term rendered its framing, not its numbers, so on them the staleness was undetectable in either direction. Claude was the only window because it is the only one of the five that reproduces the specific figures. This has a sharp consequence worth stating plainly: cited-version lag is systematically under-observable. Most corrections change text no engine renders verbatim, so most instances leave no trace at all. What we measured is not the rate at which it happens; it is one of the rare cases where it was visible.

And even the 48 days is bracketed, not exact. We probe weekly, so we know the correction had propagated somewhere between day 41 (still stale) and day 48 (clean), not the hour it flipped. Read it as "about seven weeks on this engine," a lower-resolution measurement, not a precise SLA.

The mechanism, honestly

We cannot see the internal cause from a consumer interface, so we name the behavior, not the mechanism. What we observed is consistent with an engine serving a stored, stale representation of a page whose live URL it cites correctly. That such a thing exists is documented: an independent investigation of one major engine's retrieval stack found a shared reading cache that serves stored page copies, some observed more than ninety days after they were fetched, and that engine's own documentation notes that in one mode its indexed or cached pages can be older than the live version. Whether what we saw was that cache, a search-index lag, a re-crawl gap, or something else is not identifiable from the answers alone. It is a timing artifact of retrieval, not fabrication (the engine cited the right URL and reproduced content the page genuinely used to carry), and it is different from a knowledge cutoff (the engine was browsing, not answering from training data).

What to do if you correct a specific fact

  • Treat a correction as submitted, not landed. After you fix a specific figure, re-query that exact figure on the engines that render specifics. An answer that only reproduces your framing is not evidence the fix propagated; it is evidence the changed part was never shown.
  • Re-probe, do not re-edit. The page is already right. Editing it again does not force a re-crawl, and it churns a correct page. Log the lag and wait.
  • Expect a timescale of weeks, and that you cannot force it. In this one case it was about seven weeks, and it resolved on its own. The publisher can invite a refresh (sitemaps, a substantive diff) but cannot set its clock.

Honest limits

This is one correction, on one engine, on one page, tracked at weekly resolution, on a site we both edit and measure. The shape (persist for several weeks, flicker, then self-heal) is what we would carry forward; the specific 48 days is not a rate and should not be quoted as one. The method is the durable part: a versioned correction timestamp paired with a weekly directed re-probe is a measurement only a party that makes dated corrections and probes longitudinally can produce, which is also why there is so little public data on how long corrections actually take to reach AI search. This is the first case; when a second number-changing correction comes through our own pages, we will run the same protocol and report the range across cases. Until then: one measured instance, honestly bounded, of a failure mode that is real, self-healing, and mostly invisible.

More dispatches