Your AI mention rides on your own citation
Crawlmind Engineering··5 min read
An AI brand mention is the engine naming your company in the answer text, and new data suggests that for most brands it depends on whether the engine retrieved a page from your own domain in that same run.
That sounds obvious until you look at how teams report AI visibility. Mentions and citations usually sit in two separate columns, tracked as if they were independent outcomes with independent levers. A paper posted to arXiv on September 19, 2026 argues they are closer to one pipeline, with the citation as an early stage and the mention as a late one. A Semrush study from June shows the other half of the picture: being cited is far from a guarantee of being named.
#What the stage model measured
Benjamin Tannenbaum's "From Prompt to Recommendation" fits a stage model of brand visibility on monitoring data from 75 anonymized projects and 2,854 distinct prompts, collected on GPT and Gemini between June and September 2026. The larger validation panel holds 34,960 unbranded prompt-engine observations. Unbranded matters here: these are prompts that do not contain the brand name, which is where discovery happens.
The paper defines two signals. Own-domain exposure means at least one source URL in the answer belongs to the target organization's domain. A branded fan-out means the organization's name shows up in at least one of the search queries the engine issued while building the answer.
The mention rates split sharply on those two signals:
- With no own-domain exposure and no branded fan-out, the brand was mentioned in 2.8% of GPT answers and 3.8% of Gemini answers (Tannenbaum, 2026).
- With own-domain citation alone, the rate rose to 49.0% on GPT and 58.4% on Gemini (Tannenbaum, 2026).
- With both a citation and a branded fan-out, it reached 91.4% on GPT and 100% on Gemini (Tannenbaum, 2026).
Exposure itself was uncommon. Only 10.8% of the unbranded GPT observations and 20.9% of the Gemini observations had any own-domain citation at all (Tannenbaum, 2026). So for most prompts in the panel, the brand was sitting in the low-single-digit bucket.
#The citation-free path belongs to incumbents
There is a second route to a mention, and the paper is careful to report it. Brands that the model already associates with a category get named without any citation. In one case the paper describes, a competitor appeared in 74 of 160 repeated ChatGPT runs while receiving only 7 citations, and in 40 runs with no citations at all it still appeared 30 times (Tannenbaum, 2026).
Prior visibility is the best single predictor in the model. A brand that was not mentioned in the previous run and had no exposure in the current one was mentioned 1.6% of the time on GPT and 1.9% on Gemini. A brand mentioned previously and exposed again was mentioned 80.5% and 83.7% of the time. Prior history alone gave a holdout AUC of 0.937 on GPT, against 0.880 for live signals alone and 0.963 for the full model (Tannenbaum, 2026).
Read together, those numbers describe two populations. Established names can be recommended from what the model already believes. Everyone else has to be read in the moment, and for them the citation is close to a precondition.
#A citation is necessary, not sufficient
The Semrush "ghost citations" study, run with Kevin Indig and Growth Memo, looked at the same relationship from the citation side. Across 3,981 domain appearances from 115 prompts in 14 countries, 61.7% were citations with no brand name in the answer, 13.2% were both cited and mentioned, and 25.1% were mentions without a citation (Semrush, June 2026).
Engines differ a lot. ChatGPT cited in 87% of appearances but named the brand in 20.7%. Gemini ran the other way, with an 83.7% mention rate and a 21.4% citation rate (Semrush). Query type mattered too: informational prompts produced an 18% mention rate against 43.3% for comparative prompts.
This does not contradict the stage model. The Semrush sample counts every cited domain, including publishers, forums and reference sites that are cited for facts and never recommended as a vendor. The stage model conditions on a target brand in unbranded prompts. Both point to the same practical reading: getting your page into the source list moves you from almost no chance to roughly a coin flip, and something else decides the rest.
#Where the stage model says to spend effort
The paper also measures page match, meaning how well the best page on the domain fits the prompt. On Gemini, the top quartile of page match saw own-domain citation 50% of the time against 22% in the bottom quartile, a risk ratio of 2.27 (Tannenbaum, 2026). The author notes that this is much smaller than the gap between exposed and unexposed runs, and recommends separating content work from distribution work. If no page matches the prompt, write one. If a page matches and still is not retrieved, the problem is discoverability, not copy.
That split maps onto a few concrete changes to how you measure:
- Report mention rate conditional on citation. Break your mention rate into two numbers: the share of runs where your domain was cited, and the mention rate within those runs. A low first number is a retrieval problem. A low second number is a positioning problem, and the Semrush data suggests it is common even for sites that are cited heavily.
- Log fan-out queries. Where an engine exposes its search queries, record whether your brand name appears in them. In the paper, a branded fan-out was the difference between about half and nearly all of exposed runs producing a mention.
- Separate incumbents from challengers in competitor reports. A rival that is named without citations is running on prior visibility. You will not displace it by matching one page. You close that gap by being retrieved consistently over time, since the model's own history is what carries mentions forward.
- Use prompts that do not contain your name. Branded prompts mostly test whether the engine can find a company it was told about. Discovery is measured on unbranded ones.
#Caveats worth keeping
The paper is observational. The author states the equation is "predictive and observational, not a causal description" of engine internals, and that the datasets were collected for operational research rather than as one preregistered experiment (Tannenbaum, 2026). The engines tested include lighter variants, GPT-5 nano alongside GPT-5.4 and Gemini 2.5 Flash, which may not behave like the consumer products your buyers use. The author is also the founder of Aiso, a company that sells AI-search measurement software, and declares that financial conflict of interest in the paper.
The Semrush sample is small by comparison, 115 prompts, and was drawn from its own AI Visibility Toolkit. Neither study tells you how your category behaves. Both give you a structure to check it against: count citations, count mentions, and look at the rate that connects them. If the mention only rarely arrives without the citation, the citation is the metric to move first.
Related field notes
September 30, 2026 · 5 min
Brand mentions follow the retrieval path
New data on 34,960 unbranded prompts: when an AI engine never retrieves your domain or searches your name, it almost never recommends you.
September 30, 2026 · 6 min
Find the stage where your brand drops out
A new 34,960-observation study splits AI brand visibility into match, exposure, fan-out and selection. Each failing stage needs a different fix.
September 28, 2026 · 5 min
The AI Overview is now a door to AI Mode
Google keeps adding paths from AI Overviews into AI Mode. Each one moves searchers into the surface that sends the fewest clicks, and Search Console blends both.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.