We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

AI engines reward reach, not authority scores

Crawlmind Engineering··5 min read

Citation reach is how often an AI engine cites a website on prompts outside the topic you care about, and a new study finds it predicts that site's citations on fresh prompts better than the 0-100 domain authority scores that still appear in many SEO reports.

The study comes from Wellows, an AI visibility vendor, and it was announced on October 9, 2026. It is observational, the authors say so plainly, and it has real limits. It is still the cleanest public test so far of a question every agency gets asked: does a high authority score mean AI engines will cite you?

#How the study was built

Wellows used 9,471 English-language questions across 151 topics and 247,397 answer snapshots collected from January to May 2026, covering ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode. The press release gives slightly different totals (382,176 AI answers, January to June), so treat the research page as the reference for method. About 86% of collection was in the United States.

The design choice that matters is the split. Within each topic, questions were divided into a measurement half and a held-out half. Each site's measures were calculated on the first half, then tested against which sites got cited on the second. That avoids the circular result you get when a "visibility score" is built from the same answers it claims to predict.

Three measures were compared, each with a within-topic Spearman rank correlation averaged across the 151 topics:

#What it found

Reach came first on every engine. Its correlation with held-out citations ran from 0.29 on ChatGPT and Perplexity to 0.37 on Gemini, with Google AI Mode at 0.35 and AI Overviews at 0.33.

Coverage beat the authority score on four of five engines. The figures were 0.29 against 0.22 on AI Overviews, 0.28 against 0.22 on AI Mode, 0.25 against 0.15 on Perplexity and 0.13 against 0.10 on ChatGPT. Gemini was the exception, with authority at 0.21 and coverage at 0.20.

The most useful table compares sites with similar reach. With reach held roughly constant, coverage still carried signal on Google's surfaces: 0.208 on AI Overviews and 0.197 on AI Mode, against shuffled baselines near 0.01. On ChatGPT the figure falls to 0.043 against a 0.012 baseline. ChatGPT appears to care about who you already are across the web more than about whether other engines trust you on this specific topic.

None of these correlations are large. A coefficient of 0.3 leaves most of the variation unexplained. The ranking between measures is the finding, not the size of any one number.

#Why an authority score falls behind

Domain authority style metrics were built to estimate how a site might rank in classic search, mostly from link graphs. Google has said for years that it does not use them. John Mueller's line, as collected by Search Engine Journal, was "We don't use domain authority at all in our algorithms." For AI Overviews and AI Mode, Google's own documentation says a page must be indexed and eligible to be shown with a snippet, with no additional requirements.

So an authority score is at best a proxy for a proxy. Reach and coverage are different in kind. They are measured on the engines themselves, from their own citation behavior. A measure taken from the outcome predicting the outcome better than a measure taken from link counts is the expected result. The study's value is putting numbers on the size of that gap, per engine.

#What the study does not show

The authors are explicit about three limits, and each one changes how you should act on the result.

First, it is an association. The study states that it did not test whether publishing more changes it. Wide reach could reflect brand recognition, editorial history or press coverage that no content calendar can reproduce in a quarter.

Second, the findings cover sites already cited in the measurement questions. They say nothing about how an uncited site earns its first citation, which is the problem most smaller brands actually have.

Third, the study reports no confidence intervals or repeated permutation tests, so the gaps between engines should not be read as established differences. The ChatGPT result is suggestive, not proven. It also comes from one vendor's prompt set, mostly in the US.

#What to change in your reporting

Stop leading with an authority score. Keep it as context for link work if you like, but do not present it as an AI visibility metric. On this data it was the weakest of the three predictors on four engines.

Measure reach directly. Track how often each engine cites your domain across a broad, stable prompt set, not only on the prompts in your core category. A narrow tracked set will miss the signal that predicted citations best here.

Read coverage across engines. If AI Overviews and Perplexity cite you on a topic and ChatGPT does not, that gap is information. Coverage carried weight on Google's surfaces even among sites with similar reach. On ChatGPT it barely did, which points to broader brand presence as the lever there.

Split your tracked prompts before you judge a tactic. The held-out design is easy to copy. Set aside part of your prompt set, make the change, and check whether citations move on prompts you did not tune for. If they only move on the prompts you optimized against, you have measured the tuning, not the site.

Separate the cold-start problem. If your site is not cited at all, this study does not describe your situation. The first citation usually depends on being indexed, being retrievable for a specific question, and being named by third-party sources the engine already trusts. Work on those before you worry about reach.

Keep engines apart. The same measure behaved differently on ChatGPT and on Google's surfaces. One blended "AI visibility" number would hide that.

The practical reading is narrow but useful. Engines appear to favor sites they already cite widely, and they partly follow each other on a topic. A link-based authority score adds little on top. Measure the thing the engines do, on prompts you did not optimize for, and report each engine on its own line.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.