We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

AI engines agree on brands, not on URLs

Crawlmind Engineering··4 min read

Cross-engine citation overlap is the share of sources that two or more AI answer engines cite for the same question, and it is not one number. It moves by roughly a factor of five depending on whether you count exact URLs, domains, or brand names in the answer text.

That distinction decides strategy. Read the URL-level figure and the reasonable conclusion is that ChatGPT, Gemini, Perplexity and Google's two answer surfaces are separate ecosystems needing separate playbooks. Read the brand-level figure from the same dataset and the conclusion is much milder. Both readings are in circulation right now, usually without anyone saying which layer produced them.

#One dataset, three answers

Wellows analyzed 22.7 million citations across 1.15 million questions between January and June 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode, with 84% of the sample in the United States. For the subset where all five engines returned sources, the study reports agreement at three different resolutions.

At the page level, meaning the exact URL, engines agreed on 6.8% of citations. At the website level, 79.6% of the sites cited on a question appeared on only one of the five engines, and all five converged on the same source 0.31% of the time. Measured on brand names appearing in the answer text, agreement reached 30.3%.

Same crawl, same questions, three numbers that support three different strategy memos.

The pairwise splits are just as stark. Of the websites ChatGPT cited, Perplexity never touched 89.1% on the same question, and read in reverse, ChatGPT missed 90.1% of Perplexity's.

#A second dataset with the same shape

Writesonic ran a narrower version of the experiment and landed in the same place: 161,286 prompts across ChatGPT, Gemini, Perplexity and AI Overviews in May and June 2026, of which 70,879 returned citations from all four engines. All four agreed on 3.8% of sources, and 72 to 73% of every cited domain appeared on exactly one engine.

The methodology footnote is the whole point. Writesonic measured at the domain level, lowercasing URLs and stripping www, and states plainly that exact-URL matching would produce lower scores. Its most similar pair, Perplexity and AI Overviews, scored 0.237 on Jaccard similarity; its least similar, ChatGPT and Gemini, 0.119. Those sit above the Wellows page-level figure and below its website-level one, which is exactly where a domain-level measure should land.

Two independent datasets, consistent ordering. The disagreement in the market is not about the data.

#Google is not one engine either

Fragmentation usually gets framed as a between-vendor problem. It is not. Profound tracked 15,155 brand configurations daily through May 2026 across Google's own three surfaces and found the median brand facing a daily visibility gap of 8 percentage points between its best and worst Google model, with more than one in three brands seeing that gap exceed 10 points.

At the entity level, AI Overviews and AI Mode shared about 40% of their company mentions, while Gemini shared 27% with AI Overviews and 29% with AI Mode. The baseline in that same study is the number worth holding onto: a single model compared against itself on alternating days shares just over half of its company mentions. Run-to-run noise inside one model is the floor, and the gap between two different Google models is not dramatically wider than it.

Citation volume differs too, which mechanically drags overlap down. Gemini returned 6.6 citations per run, AI Overviews 11.1, and AI Mode 15.2. A short answer and a long one cannot overlap much even when they agree completely on which sources rank highest.

#Why the layers come apart

The gap between URL agreement and brand agreement is not a rounding artifact. The two are structurally decoupled, because the page an engine cites is usually not a page belonging to the brand it names.

Shero Commerce examined 1,851 cited sources from buying prompts across Google AI Mode, ChatGPT and Perplexity and found that only 2.8% were brand-owned pages, with third-party sites accounting for 59%. When a brand was actually recommended by name, its own page was cited in 31% of those cases.

The two layers are measuring different events. One asks whether the answer names you. The other asks whether the answer links you. Engines converge on the first far more than on the second, and they source that convergence from review sites, retailers and forums they do not agree on either.

#Which layer to count

Pick the layer that matches what you are trying to change.

If the objective is referral traffic, count URLs. That is the layer that produces clicks, it is the layer with the least cross-engine agreement, and per-engine work is genuinely justified there.

If the objective is being recommended inside the answer, count brand mentions. Fragmentation at this layer is real but much smaller, and the work that moves it (third-party coverage, review presence, a consistent entity description) is largely shared across engines. Building five separate content strategies on the strength of a URL-level statistic is the common and expensive error.

If you are comparing your own dashboard to a published study, check the layer before you argue with the number. A vendor reporting domain-level overlap and one reporting page-level overlap will differ by a multiple with no methodological fault on either side. Ask which resolution, which engines, and what share of prompts returned citations from all of them, because that last filter changes the denominator more than most people expect.

One more thing worth watching. Wellows reports monthly average agreement rising from 8.49% in January to 11.25% in June 2026, about a third higher over six months. If that direction holds, the per-engine playbook gets less valuable over time, not more. Treat today's fragmentation figure as a snapshot with a trend attached rather than a fixed property of AI search.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.