Three things happen before an answer mentions your product
Fetching. A crawler requests URLs. It does not execute your storefront's JavaScript, so anything rendered in the browser is invisible to it. Between 25.07.2026 and 20.09.2026 we recorded 16,141 crawler requests to our own product records and logged, for each one, whether the visitor could prove it was who it claimed to be.
Reconstruction. A model rarely answers from one page. It assembles a product from a manufacturer page, a retailer listing, a marketplace card, a review, sometimes a PDF spec sheet. Those sources disagree about capacity, model number, what is included in the box, and which variant is which. The model resolves the disagreement silently.
Answering. What reaches the buyer is the reconstruction, with a destination attached. The destination is frequently not the manufacturer.
Not every "AI crawler" is an AI crawler
A user agent string is a claim, not an identity. Anyone can send GPTBot in a header. The providers publish either IP ranges or reverse DNS records so that a claim can be checked, and we check every request against them. Our own log, 25.07.2026 to 20.09.2026:
| Provider | Passed the check | Failed it | Only said so in the header | | --- | --- | --- | --- | | Google | 2,337 | — | — | | OpenAI | 1,303 | 2,290 | — | | Perplexity | 518 | 715 | — | | Apple | 911 | — | — | | Anthropic | — | — | 2,487 | | Meta | — | — | 1,633 | | Amazon | — | 2,899 | — |
Read this carefully, because the obvious reading is wrong. "Failed" means the request did not match the published ranges at the moment we checked — it does not prove impersonation, and a provider our check has nothing published to compare with can only be recorded as "said so in the header". What the table does establish: if you count AI traffic by user agent alone, a large share of what you count is unverifiable.
What models cite when they answer about products
Since 26.07.2026 we have run 4,946 probes against five providers and stored 4,078 answers with 22,789 extracted citations. The models and how many days each has been measured, named in full because behaviour changes under a stable name: gpt-4o-mini (OpenAI) 39 days, sonar (Perplexity) 30, grok-4-latest (xAI) 18, gemini-2.5-flash (Google) 15, claude-haiku-4-5-20251001 (Anthropic) 4, mistral-small-latest (Mistral) 1.
Two observations hold across that period. Answers cite pages, not files: the machine-readable formats we publish are fetched far less than the HTML record that carries the same facts. And a citation is not a fact: a model can link a page and state something the page does not say.
What you can actually control
- Identity. GTIN, MPN, brand, and the variant axis, stated in the page
itself rather than implied by the title. Without them a model cannot tell your 45-litre pack from a retailer's listing of last year's 48-litre one.
- Reachability without JavaScript. If the facts appear only after the
bundle runs, they are not in what the crawler read.
- An official destination. If your page is not a plausible thing to link,
the retailer's page is.
- Provenance. Which source said what, and when it was read. A fact with a
source survives a disagreement; a fact without one loses to whichever page the model found first.
What none of this proves
Publishing a record does not make a model cite it. Being crawled does not mean being retrieved. Being cited does not mean the cited fact was used. We keep those as separate, separately recorded levels — the ladder is at /evidence — and no page on this site will tell you that publishing causes a recommendation. What we can show is which step of the ladder your product has reached, with dates.
Sources
- Evidence ladder and what each level does not prove
- Published product records
- Machine-readable entry points we publish