- Period
- 2026-08-04 — 2026-09-16
- Method
- Each recorded answer carries an outcome: whether the product was matched and whether an official URL was present, graded strict, partial, candidate or text. Answers are grouped by product and question; a group counts as comparable when at least three models answered it.
- Method version
- 1.0
- Published
- 2026-09-20
What was measured
2 185 graded answers between 2026-08-04 and 2026-09-16, each one a real buyer question about a real product. The grade records what the answer did for the seller: whether the product was matched, whether an official URL was present, and where that URL pointed.
Agreement
54 questions were put to three or more models. Of those:
- 47 (87%) produced the same
outcome from every model, of which 33 were unanimously the strict outcome — product matched and official URL present;
- 7 (12%) split: at least one model reached
an outcome the others did not.
Read this figure narrowly. It compares answers that already produced a recorded outcome, so it says that models which do answer usefully tend to agree about what they found — not that models agree in general. The disagreement that matters to a brand is one step earlier and much larger: whether a model produces a usable answer at all, which the table below shows ranging from 52% to 83%.
Per model
| Provider | Model | Completed runs | Graded answers | Share that produced one | Strict, of graded | | --- | --- | --- | --- | --- | --- | | openai | gpt-4o-mini | 1 970 | 1 047 | 53% | 56% | | perplexity | sonar | 1 393 | 903 | 64% | 83% | | anthropic | claude-haiku-4-5-20251001 | 156 | 89 | 57% | 73% | | xai | grok-4-latest | 159 | 84 | 52% | 64% | | gemini | gemini-2.5-flash | 72 | 60 | 83% | 66% |
"Graded answers" counts answers that produced a recorded outcome at all. An answer that named no product and carried no official URL leaves no row, which is why the third column matters more than the fourth: it is the share of questions where the seller got anything.
Where the link actually points
When an official URL is present it is not always the product's own page:
- the product's page: 643 (29%);
- the right site, another page: 1 445 (66%);
- the bare host, no page: 97 (4%).
This distinction is usually collapsed in reporting on AI visibility, and it is the difference between a buyer landing on the thing they asked about and a buyer landing on a home page with a search box.
What this does not say
Nothing here explains why a model answered as it did, and nothing here is a ranking. The grades describe outcomes we can check in the text of an answer: whether the product was named and whether the seller's URL was present. The record does not show what the model read, and a difference between models is not evidence about the quality of either.
Questions put to fewer than three models are excluded from the agreement figures; the per-model table uses every graded answer, so the two sections count different things on purpose.