We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

Brand mentions follow the retrieval path

Crawlmind Engineering··5 min read

A retrieval path is the chain of background searches and sources an AI engine pulls in between a user's prompt and its answer, and new data suggests that a brand missing from that chain is almost never the one recommended.

That sounds obvious. The size of the gap is not.

#The ladder

On September 19, 2026, Benjamin Tannenbaum posted From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search. The main panel covers 34,960 unbranded prompt observations across 75 projects and 2,854 monitored prompts, collected from June to September 2026 on GPT and Gemini.

Each observation is sorted by two signals the engine exposes in its own trace. Own-domain citation means the brand's registered domain appears in the stored source list. Branded fan-out means the brand name appears in at least one of the background search queries the engine issued. Google describes that fan-out step for AI Mode as "breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf". The paper reads the equivalent search queries from the GPT trace as well.

Here are the brand mention rates for each cell, from the paper's full text:

Evidence in the trace GPT mention rate (n) Gemini mention rate (n)
Neither signal 2.8% (15,524) 3.8% (13,801)
Own domain cited only 49.0% (1,769) 58.4% (3,415)
Branded fan-out only 64.4% (59) 84.8% (33)
Both signals 91.4% (128) 100.0% (231)

The paper puts the own-domain step at a 17.4-fold increase in mention probability for GPT and 15.4-fold for Gemini over the neither-signal baseline. Look at the sample sizes before quoting the lower rows. The fan-out-only cells hold 59 and 33 observations. The top row holds most of the data, and it says that without evidence in the path, a brand shows up in roughly one answer in thirty.

#What the ladder does not prove

The author is direct about the limits, and so should anyone citing this. The paper calls the result "an observational association, not a randomized causal effect". Causation can run both ways here. An engine that already leans toward a brand may search for it by name, which is exactly what the branded fan-out signal records.

Prior history also predicts most of the outcome on its own. On a 30% holdout, a model using only a brand's past mention rate for the same prompt scored an AUC of 0.937 for GPT and 0.917 for Gemini. Live retrieval signals alone scored 0.880 and 0.840. The full model reached 0.963 and 0.942. Past visibility is the best single predictor of future visibility, which is uncomfortable for anyone trying to break in.

Two more caveats. Own-domain citation excludes third-party pages, which the author says "can be more important than owned pages". The author also discloses a conflict: the author founded and runs a company that sells AI search measurement software.

#Challengers live or die on exposure

The most useful table in the paper splits results by prior. For low-prior brands (historical mention rate below 0.1), GPT mentioned the brand 0.9% of the time without exposure and 25.1% with it. Gemini went from 1.0% to 23.8%. For high-prior brands (0.5 or above), GPT went from 32.0% to 82.0%.

Read those two rows side by side. An established brand keeps a floor of about one mention in three even when nothing about it is retrieved. A challenger has no floor. For the challenger, retrieval is the whole game. This matches what we wrote about incumbency breaking on evidence: the prior wins only when the retrieved material gives the model nothing to separate the options.

A small repeated-run case in the paper makes the same point. Across 160 runs of 16 prompts, the author isolated 40 runs in which the engine returned zero citations. Two competitors were mentioned in 75% and 70% of those runs. The studied organization was mentioned in none. With no retrieval, only memory is left, and memory favors whoever was already well known.

There is a partial reprieve for commercial queries. Commercial prompts triggered fan-out in 78.3% of cases against 3.6% for informational prompts. When someone asks which product to buy, the engine usually goes looking, so the retrieval path is open to challengers exactly where it matters most.

#Off-site evidence feeds the same path

The paper measures only the brand's own domain, but the path also carries review sites, forums, and publisher pages. An earlier Ahrefs analysis of 75,000 brands found that branded web mentions correlated with AI Overview visibility at 0.664, against 0.218 for backlinks. That is also a correlation, and it measures a different surface. Taken together, the two studies say something consistent: the brand name has to be present somewhere the engine reads at answer time.

#What to change

Make your domain retrievable first. OpenAI's crawler documentation states that sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers". A blanket bot block, or a CDN rule that challenges search crawlers, puts you in the neither-signal row by default.

Map prompts to pages. The author's advice is to start with "which real requests is this page supposed to satisfy?" If a page matches the request well and still is not retrieved, publishing more on the same domain may not help. The gap is then off-site: comparisons, reviews, and third-party coverage that name you.

Log the trace, not only the answer. Record the fan-out queries and the source list on every run. A citation dashboard and a mention dashboard answer different questions, and the paper recommends reporting them separately by engine.

Separate citation-free runs. Answers with no retrieval measure the model's prior, not your content. Mixing them into one visibility score hides whether a change moved retrieval or nothing at all. If your share looks flat, check how many reruns you are averaging before concluding anything.

The practical reading is narrow and useful. You cannot edit an engine's memory of your category this quarter. You can decide whether your pages and your name are present when it searches.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.