AI engines reward reach, not authority scores
Crawlmind Engineering··5 min read
Citation reach is how often an AI engine cites a website on prompts outside the topic you care about, and a new study finds it predicts that site's citations on fresh prompts better than the 0-100 domain authority scores that still appear in many SEO reports.
The study comes from Wellows, an AI visibility vendor, and it was announced on October 9, 2026. It is observational, the authors say so plainly, and it has real limits. It is still the cleanest public test so far of a question every agency gets asked: does a high authority score mean AI engines will cite you?
#How the study was built
Wellows used 9,471 English-language questions across 151 topics and 247,397 answer snapshots collected from January to May 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode. The press release gives slightly different totals (382,176 AI answers, January to June), so treat the research page as the reference for method. About 86% of collection was in the United States.
The design choice that matters is the split. Within each topic, questions were divided into a measurement half and a held-out half. Each site's measures were calculated on the first half, then tested against which sites got cited on the second. That avoids the circular result you get when a "visibility score" is built from the same answers it claims to predict.
Three measures were compared, each with a within-topic Spearman rank correlation averaged across the 151 topics:
- Outside-topic reach: how often the site is cited on prompts about other topics.
- Citation coverage: the share of measurement snapshots in which at least one of the other four engines cites the site. It is a cross-engine signal, not a count of pages you published.
- Authority: the 0-100 domain authority score Wellows stores with each citation, summarized as each site's approximate median.
#What it found
Reach came first on every engine. Its correlation with held-out citations ran from 0.29 on ChatGPT and Perplexity to 0.37 on Gemini, with Google AI Mode at 0.35 and AI Overviews at 0.33.
Coverage beat the authority score on four of five engines. The figures were 0.29 against 0.22 on AI Overviews, 0.28 against 0.22 on AI Mode, 0.25 against 0.15 on Perplexity and 0.13 against 0.10 on ChatGPT. Gemini was the exception, with authority at 0.21 and coverage at 0.20.
The most useful table compares sites with similar reach. With reach held roughly constant, coverage still carried signal on Google's surfaces: 0.208 on AI Overviews and 0.197 on AI Mode, against shuffled baselines near 0.01. On ChatGPT the figure falls to 0.043 against a 0.012 baseline. ChatGPT appears to care about who you already are across the web more than about whether other engines trust you on this specific topic.
None of these correlations are large. A coefficient of 0.3 leaves most of the variation unexplained. The ranking between measures is the finding, not the size of any one number.
#Why an authority score falls behind
Domain authority style metrics were built to estimate how a site might rank in classic search, mostly from link graphs. Google has said for years that it does not use them. John Mueller's line, as collected by Search Engine Journal, was "We don't use domain authority at all in our algorithms." For AI Overviews and AI Mode, Google's own documentation says a page must be indexed and eligible to be shown with a snippet, with no additional requirements.
So an authority score is at best a proxy for a proxy. Reach and coverage are different in kind. They are measured on the engines themselves, from their own citation behavior. A measure taken from the outcome predicting the outcome better than a measure taken from link counts is the expected result. The study's value is putting numbers on the size of that gap, per engine.
#What the study does not show
The authors are explicit about three limits, and each one changes how you should act on the result.
First, it is an association. The study states that it did not test whether publishing more changes it. Wide reach could reflect brand recognition, editorial history or press coverage that no content calendar can reproduce in a quarter.
Second, the findings cover sites already cited in the measurement questions. They say nothing about how an uncited site earns its first citation, which is the problem most smaller brands actually have.
Third, the study reports no confidence intervals or repeated permutation tests, so the gaps between engines should not be read as established differences. The ChatGPT result is suggestive, not proven. It also comes from one vendor's prompt set, mostly in the US.
#What to change in your reporting
Stop leading with an authority score. Keep it as context for link work if you like, but do not present it as an AI visibility metric. On this data it was the weakest of the three predictors on four engines.
Measure reach directly. Track how often each engine cites your domain across a broad, stable prompt set, not only on the prompts in your core category. A narrow tracked set will miss the signal that predicted citations best here.
Read coverage across engines. If AI Overviews and Perplexity cite you on a topic and ChatGPT does not, that gap is information. Coverage carried weight on Google's surfaces even among sites with similar reach. On ChatGPT it barely did, which points to broader brand presence as the lever there.
Split your tracked prompts before you judge a tactic. The held-out design is easy to copy. Set aside part of your prompt set, make the change, and check whether citations move on prompts you did not tune for. If they only move on the prompts you optimized against, you have measured the tuning, not the site.
Separate the cold-start problem. If your site is not cited at all, this study does not describe your situation. The first citation usually depends on being indexed, being retrievable for a specific question, and being named by third-party sources the engine already trusts. Work on those before you worry about reach.
Keep engines apart. The same measure behaved differently on ChatGPT and on Google's surfaces. One blended "AI visibility" number would hide that.
The practical reading is narrow but useful. Engines appear to favor sites they already cite widely, and they partly follow each other on a topic. A link-based authority score adds little on top. Measure the thing the engines do, on prompts you did not optimize for, and report each engine on its own line.
Related field notes
October 10, 2026 · 5 min
Google now names fake bylines as deception
Google's helpful content guide now calls invented authors, AI headshots and false credentials deception. AI answers inherit that judgment.
October 9, 2026 · 6 min
Model swaps are the new core updates
When an AI engine changes its default model, citations move overnight. How to tell a model swap from a content problem, and what to log so you can.
October 6, 2026 · 5 min
AI mentions follow your own citations
A new panel study finds unbranded AI answers rarely name a brand unless its own domain is in the sources. Split content fixes from retrieval fixes.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.