Find the stage where your brand drops out
Crawlmind Engineering··6 min read
A brand visibility stage model is a way of breaking one AI answer into the separate steps that decide whether your brand appears: whether your page fits the prompt, whether the engine retrieves it, whether the engine searches for you by name, and whether the model names you once the evidence is in front of it.
Most AI visibility reporting collapses all of that into one number: the share of prompts where the brand was mentioned. A September 2026 paper by Benjamin Tannenbaum shows why that single number is hard to act on. When a brand is missing, it can be missing for four different reasons, and the fix for one does little for the others.
#What the study measured
The paper tracks GPT and Gemini answers across 34,960 unbranded observations covering 2,854 monitored prompts in 75 projects, collected between June and September 2026. Unbranded means the prompt never contained the brand name. The question is how a brand gets into an answer nobody asked it to be in.
It separates five components:
- Match: how well your best page answers the prompt, scored with BM25 against the text of your site.
- Exposure: whether evidence supporting your brand shows up in the engine's retrieval or citation set.
- Fan-out: whether the engine expands the prompt into its own search queries, and whether any of those queries contain your name.
- Selection: whether the model mentions you once that evidence is available.
- Prior: the chance the model mentions you with no visible evidence at all, from training data or sources the study could not observe.
Fan-out is a documented behavior, not a research construct. Google says AI Overviews and AI Mode may issue "multiple related searches across subtopics and data sources" to build a response. OpenAI says ChatGPT search typically rewrites a prompt into one or more targeted queries before sending them to search partners.
#The four-cell ladder
The clearest result is a simple table. It splits every observation by two signals: was your own domain cited, and did the engine run a search that contained your brand name. From the paper's large-panel results:
| Signal present | GPT mention rate | Gemini mention rate |
|---|---|---|
| Neither | 2.8% | 3.8% |
| Own domain cited only | 49.0% | 58.4% |
| Branded fan-out only | 64.4% | 84.8% |
| Both | 91.4% | 100.0% |
Two things stand out. First, with neither signal, a brand almost never appears. Second, a branded search query by the engine is a stronger signal than your own domain being cited. If the engine goes looking for you by name, you are very likely to be named.
Read this as association, not proof of cause. A brand the engine already knows well is more likely both to be searched for by name and to be mentioned. The author treats the ladder as a diagnostic, and the paper lists that limitation openly.
#Fan-out mostly happens on commercial prompts
In a smaller set of 80 real user prompts, 78.3% of commercial prompts triggered observed fan-out against 3.6% of informational ones. The 20 prompts that did trigger it produced 42 search queries, and 39 of those were labeled commercial.
That matters for how you build a prompt set. If your tracked prompts are mostly informational, you are measuring the part of AI search where the engine often answers from memory. The buying questions, the ones like "best tool for X" or "X vs Y", are where the engine goes out and searches, and where retrieval can change the outcome.
The author's advice is to start from real conversations or real prompt logs rather than a synthetic keyword list. The fan-out queries themselves are useful data: they show which criteria the engine decided to check on the user's behalf.
#Citation and mention are different outcomes
A cited domain does not guarantee a mention, and a mention does not need a citation. In one benchmark of 16 prompts run 10 times each, a competitor was mentioned 74 times while its own URLs were cited only 7 times. In the 40 runs where it had no citation at all, it was still mentioned in 30.
The reverse also happened. The benchmarked organization had a page in Bing's top 30 for three of the sixteen prompts. In one of those cases ChatGPT mentioned the brand in 0 of 10 runs while competitors appeared in 6 to 9. Ranking in the index the engine reads from did not carry the brand into the answer.
This is why a citation count and a mention count should sit in separate columns. One tells you about retrieval, the other about selection.
#History predicts the next answer
The strongest single predictor in the paper is a brand's own past mention rate. On held-out data, a model using only prior visibility scored an AUC of 0.937 for GPT and 0.917 for Gemini. Adding live evidence raised that to 0.963 and 0.942.
Exposure still moves low-prior brands. Where a brand's historical mention propensity was below 0.1, GPT mentioned it 0.9% of the time without exposure and 25.1% with it, while Gemini went from 1.0% to 23.8%. For a brand the engines do not yet know, getting into the retrieval set is the main lever available.
#Match the fix to the stage
The practical value of the model is that each failing stage points to a different kind of work. The paper states the first two directly, and the rest follow from the ladder.
- Low match: your site has no page that answers the prompt. This is content work. Write the page that answers the request directly.
- High match, low exposure: the page exists but the engine does not retrieve it. More content on the same domain may not solve that bottleneck. This is distribution work: third-party pages, reviews, comparisons and listings the engine already retrieves. The author notes third-party pages can matter more than owned ones.
- Exposure without mention: the engine sees your evidence and still names someone else. Look at what the cited page says about you, and how it compares you, rather than at whether it was fetched.
- No branded fan-out: the engine never searches for you by name. That usually tracks brand familiarity, which grows from repeated appearances in the sources the engine already uses.
The paper also argues against a single engine-agnostic score. GPT and Gemini differ at every rung of the ladder, so the safer report scores page fit once and then shows exposure and selection separately for each engine.
#Limits worth keeping in view
The author is careful about scope. The page crawl was run in September against a July visibility export, which may overstate fit. A prompt with no recorded search query was treated as having no fan-out, so hidden retrieval would weaken that reading. Some cells in the ladder are small. Only 20 of the 80 real prompts triggered fan-out. A mention is not the same as a positive recommendation, a click or a sale, and all of it is a dated measurement of systems that keep changing.
None of that removes the main point. When a report says your brand appears in some share of AI answers, ask which stage produced the misses. A missing page, a page the engine never retrieves, and a page it retrieves but ignores are three different problems, and one percentage cannot tell them apart.
Related field notes
September 28, 2026 · 5 min
The AI Overview is now a door to AI Mode
Google keeps adding paths from AI Overviews into AI Mode. Each one moves searchers into the surface that sends the fewest clicks, and Search Console blends both.
September 26, 2026 · 5 min
AI engines cite synthetic pages from the tail
Three studies find AI-written pages in AI answer citations, and the share rises the further engines reach past the ranked head of the web.
September 26, 2026 · 6 min
Your GEO lift assumes nobody else rewrites
Most GEO gains are measured with one optimized page among untouched rivals. When competitors rewrite too, the lift shrinks and some tactics go negative.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.