Side by side
| Step | What happened | Evidence needed | What it does not show | | --- | --- | --- | --- | | Published | a stable URL returns the page | the URL, a timestamp, a hash of the content | that anyone read it | | Accessible | it can be read without logging in or running scripts | a plain HTTP fetch returns the facts | that a crawler came | | Crawled | a crawler fetched it | a request that passed the provider's documented check | that it was used for any answer | | Retrieved | it was fetched for a specific question | a fetch traceable to a user query | that the answer linked it | | Cited | an answer linked it | the URL in an observed answer | that the answer used what the page says | | Fact matched | an answer repeated a value from it | a versioned, sourced value reproduced in the answer | that publishing caused the answer | | Influenced | the page changed the answer | a controlled experiment with one changed factor | โ |
The same ladder, with version and changelog, is published at /evidence.
Why they get merged
Because the lower steps are easy to count and the upper ones are not. Crawler requests arrive by the thousand; citations need someone to ask questions and store answers; fact matches need a versioned record to compare with; influence needs an experiment. A report that wants a big number takes it from the bottom of the ladder and labels it with a word from the top.
How big the gap is in practice
On our own pages, 2,743 were fetched by verified crawlers and 42 were later cited by the same provider (study). Merging the two steps would have overstated our own visibility by a factor of about sixty.
How to use the ladder
Ask of any claim about AI visibility: which step is this evidence from? A crawl count answers "crawled". A screenshot of one answer answers "cited", once, for one model on one day. Neither answers "the answer relied on our page", and nothing below an experiment answers "our change caused it".