We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

AI mentions follow your own citations

Crawlmind Engineering··5 min read

A brand mention in an unbranded AI answer is, in most cases, the last step of a chain that starts with the engine retrieving the brand's own pages, and a new panel study puts numbers on how tightly the two are linked.

The paper is "From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search" by Benjamin Tannenbaum, submitted to arXiv on September 19, 2026 (arXiv 2609.23162). The author works with Aiso, a commercial AI visibility platform, and the data comes from Aiso's monitoring pipeline. That is a vendor dataset, so read it as one strong observation set rather than settled fact. The structure of the findings is still useful for anyone deciding where to spend effort.

#What the study measured

The main panel covers 34,960 unbranded prompt and engine observations from 75 anonymized projects, built on 2,854 distinct monitored prompts, with repeated runs on GPT and Gemini between June and September 2026 (arXiv 2609.23162).

Three variables carry the analysis (paper HTML):

  • Mention: the target brand's name literally appears in the answer.
  • Own-domain exposure: the brand's registered domain appears in the engine's stored source list for that run.
  • Branded fan-out: the brand's name appears in at least one of the search queries the engine generated from the unbranded prompt.

The fan-out variable matters because both major engines expand prompts before they search. Google says AI Overviews and AI Mode may use "query fan-out" to issue multiple related searches across subtopics (Google Search Central). OpenAI says ChatGPT search typically rewrites a prompt into one or more targeted queries for its search partners, and may use location and saved memories when it does (OpenAI Help Center).

#The headline numbers

When neither the brand's domain nor its name showed up in the run's evidence, the brand was mentioned in 2.8% of GPT answers and 3.8% of Gemini answers (arXiv 2609.23162).

When the brand's own domain was among the sources but its name was not in the fan-out queries, the mention rate rose to 49.0% on GPT and 58.4% on Gemini. When both were present, it reached 91.4% on GPT and 100% on Gemini (arXiv 2609.23162).

The author also compared runs of the same prompt for the same organization on the same engine, which removes much of the "strong brands get cited and mentioned" confound. Within those cells, the common odds ratio for a mention when the domain was exposed was 15.3 on GPT and 29.7 on Gemini (paper HTML).

#History predicts almost as much as live evidence

The second finding is persistence. Brands with a low prior mention rate (below 0.1) and no exposure in the current run were mentioned in 0.9% of GPT answers and 1.0% of Gemini answers. Brands with a high prior rate (0.5 or above) and current exposure were mentioned 82.0% of the time on GPT and 86.2% on Gemini (paper HTML).

On a holdout set, prior history alone predicted mentions with an AUC of 0.937 on GPT and 0.917 on Gemini. Live retrieval signals alone scored 0.880 and 0.840. The full model reached 0.963 and 0.942 (arXiv 2609.23162).

In plain terms: if a brand was mentioned for a prompt last time, it will probably be mentioned again, and if it was not, live changes have to overcome that history. Visibility on a given prompt behaves more like a state than a coin flip.

The study also tested whether a better match between the prompt and the brand's best page predicts mentions. For one organization, the author crawled 275 pages and scored each of 199 prompts against them with BM25 (paper HTML). That match score predicted visibility only modestly: AUC 0.641 on Gemini and 0.545 on GPT, where 0.5 is chance (arXiv 2609.23162).

The author flags that this crawl happened after the visibility snapshot, so it is a retrospective proxy. Even so, the contrast with the exposure numbers is large. Having a relevant page on the site did much less than having that page actually land in the engine's source list.

#What the study does not show

The paper is explicit about its limits, and they should travel with any number quoted from it (paper HTML):

  • It is observational. The author does not claim that getting a URL into a source list causes a tenfold lift.
  • Own-domain citation is an incomplete measure of exposure. A third-party review or listicle can name the brand and drive the mention while the brand's domain is absent.
  • Verticals and engines are limited, and the effects need replication on larger, frozen cohorts.
  • A mention is not a click, a lead, or a sale.

The fan-out cohort is small too. Only 20 of 80 real prompts produced observed fan-out, and commercial requests triggered it 78.3% of the time against 3.6% for informational ones, which the author presents as a cohort result rather than a general rate (paper HTML).

#How we would use this

The practical value is a diagnostic order. The paper separates three stages: does a page match the request, is it retrieved into the evidence, and does the answer select the brand once the evidence is there. Each stage calls for a different fix.

Log sources on every run, not just mentions. A mention rate on its own cannot tell you which stage failed. Store the cited URLs per run and tag whether your own domain appears. A prompt with zero mentions and zero own-domain citations is a retrieval problem. A prompt with own-domain citations and no mention is a selection problem.

Treat retrieval as distribution work. If your pages are relevant but never appear in source lists, rewriting copy is unlikely to help much. Check indexing and crawl access first, then third-party coverage on the queries the engine actually runs. For Google's AI features, the documented baseline is that a page must be indexed and eligible for a snippet (Google Search Central).

Capture fan-out where you can see it. The queries an engine generates are a better target list than the prompt text. Commercial prompts in particular get expanded into evaluation queries ("best X for Y", "X vs Y", pricing, alternatives). Pages that answer those directly have more chances to be retrieved.

Report per prompt over time. Because prior visibility is so predictive, an average across all prompts hides the movement that matters: prompts flipping from "not mentioned" to "mentioned" and staying there. Track transitions, and expect new prompts to start slow.

Keep engines apart. The study's numbers differ between GPT and Gemini at every stage, and the author recommends measuring each engine separately rather than blending them into one score.

Count third-party mentions as exposure too. The paper's own-domain measure undercounts influence from pages you do not own. Classify cited sources by whether they name you, so earned coverage shows up in the same report as owned citations.

The short version for planning: before asking why an engine does not recommend you, check whether it ever read you. For most unbranded prompts in this dataset, that was the deciding step.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.