Research2026-08-04 — 2026-09-16

Do models agree about the same product?

Among 54 questions where at least three models produced a usable answer, 87% got the same outcome from all of them. The wider disagreement is upstream: how often a model produces a usable answer at all ranges from 52% to 83% depending on the model.

Period
2026-08-04 — 2026-09-16
Method
Each recorded answer carries an outcome: whether the product was matched and whether an official URL was present, graded strict, partial, candidate or text. Answers are grouped by product and question; a group counts as comparable when at least three models answered it.
Method version
1.0
Published
2026-09-20

What was measured

2 185 graded answers between 2026-08-04 and 2026-09-16, each one a real buyer question about a real product. The grade records what the answer did for the seller: whether the product was matched, whether an official URL was present, and where that URL pointed.

Agreement

54 questions were put to three or more models. Of those:

  • 47 (87%) produced the same

outcome from every model, of which 33 were unanimously the strict outcome — product matched and official URL present;

  • 7 (12%) split: at least one model reached

an outcome the others did not.

Read this figure narrowly. It compares answers that already produced a recorded outcome, so it says that models which do answer usefully tend to agree about what they found — not that models agree in general. The disagreement that matters to a brand is one step earlier and much larger: whether a model produces a usable answer at all, which the table below shows ranging from 52% to 83%.

Per model

| Provider | Model | Completed runs | Graded answers | Share that produced one | Strict, of graded | | --- | --- | --- | --- | --- | --- | | openai | gpt-4o-mini | 1 970 | 1 047 | 53% | 56% | | perplexity | sonar | 1 393 | 903 | 64% | 83% | | anthropic | claude-haiku-4-5-20251001 | 156 | 89 | 57% | 73% | | xai | grok-4-latest | 159 | 84 | 52% | 64% | | gemini | gemini-2.5-flash | 72 | 60 | 83% | 66% |

"Graded answers" counts answers that produced a recorded outcome at all. An answer that named no product and carried no official URL leaves no row, which is why the third column matters more than the fourth: it is the share of questions where the seller got anything.

Where the link actually points

When an official URL is present it is not always the product's own page:

  • the product's page: 643 (29%);
  • the right site, another page: 1 445 (66%);
  • the bare host, no page: 97 (4%).

This distinction is usually collapsed in reporting on AI visibility, and it is the difference between a buyer landing on the thing they asked about and a buyer landing on a home page with a search box.

What this does not say

Nothing here explains why a model answered as it did, and nothing here is a ranking. The grades describe outcomes we can check in the text of an answer: whether the product was named and whether the seller's URL was present. The record does not show what the model read, and a difference between models is not evidence about the quality of either.

Questions put to fewer than three models are excluded from the agreement figures; the per-model table uses every graded answer, so the two sections count different things on purpose.

Sources

Related

← All studies