Your sentiment chart is mostly noise
Crawlmind Engineering··6 min read
Sentiment in AI visibility tracking is a label describing how an engine frames your brand inside an answer, and it is the least stable number on your dashboard by a wide margin. That instability is not a vendor defect. It is a property of the thing being measured, and it sits next to a mention metric that behaves in almost the opposite way. Reporting both at the same cadence, with the same thresholds, is how teams end up explaining a tone dip that never happened.
#Mention behaves more like a switch than a dial
A June 2026 study measured brand visibility across five AI search engines over 102 brands, 3,508 completed tracking runs and 102,025 prompt responses collected between March and May 2026, producing 15,815 brand mentions and 149,912 source citations.
Its headline result is a tier gradient. Unbranded visibility ran at 72.9% for global household names, 43.6% for established mid-market and regional brands, and 11.4% for niche and small brands, a step down of 29.3 then 32.2 percentage points. Engine choice moves the level too: day-one unbranded recognition came in at 22.1% on ChatGPT, 18.7% Gemini, 23.9% Perplexity, 51.5% Claude and 12.0% Grok.
The more useful finding is buried further down. Looking at each brand-by-prompt cell, 77.5% were strictly always-mentioned or never-mentioned, and 93.2% fell outside the volatile 30 to 70 percent band. For most of your prompt set, the outcome is effectively fixed. Running the same prompt a sixth time confirms what the first run told you.
That has a direct consequence for how a tracking report should be laid out. A line chart of overall mention rate averages a large block of settled cells with a small block of live ones, and the settled block dominates the average while telling you nothing new. The live cells are the working set, and they are a minority you can name.
#Sentiment flips at 45.5%
In that same dataset, the sentiment attached to a mention flipped between runs 45.5% of the time, against 6.8% for whether the brand was mentioned at all. The authors describe sentiment as roughly 6.7 times noisier than the mention signal.
Put a metric that changes state on nearly half of observations onto a weekly chart over a few dozen prompts, and you get movement every week whatever the underlying reality. The chart will have peaks and troughs. Someone will ask what caused the dip in week three. In most cases the honest answer is resampling.
This is not an argument for ignoring how AI engines describe you. Framing is the part of AI visibility with the most direct commercial consequence, because a mention that positions you as the expensive option does different work than one that positions you as the safe default. The argument is that a metric this volatile needs a different statistical treatment than the one sitting next to it, not the same weekly delta with a red or green arrow.
#Some of the flip is your instrument, not the engine
There is a second source of noise that rarely gets separated out: the sentiment label itself is usually produced by another model. A large evaluation of LLM-as-a-judge reliability covering 21 judge models from 9 providers across 3 benchmarks, about 541,000 judgments over 118 evaluation runs, found that reproducibility and correctness are different properties.
Self-consistency was high. Test-retest reliability across independent runs reached Krippendorff's alpha of 0.889 to 0.992. Agreement with human judgment was much weaker once corrected for chance: on MT-Bench the gap between raw exact match and Cohen's kappa ran 33.8 to 41.3 percentage points across all 21 models, with the strongest judge posting 84.9% exact match but a kappa of 0.511. Simply swapping the order of two responses flipped the verdict on 3.5% to 18.9% of items. Two models combined test-retest reliability above 0.95 with position bias above 0.10, which the authors call a consistency and bias paradox.
Those are pairwise preference judgments on benchmark answers, not brand sentiment labels, so treat the mechanism as transferable and the magnitudes as not. The mechanism is the point: a judge that returns the same label every time is not thereby returning the right one. If your sentiment column comes from a classifier nobody has validated against human reads of the same answers, part of that 45.5% flip rate belongs to the classifier rather than to the engines.
#The source layer is noisier than the brand layer
A separate measurement study ran four Swiss-German campaigns across ChatGPT, Gemini, Google AI Mode and Perplexity over 45 days with 8 prompts per campaign, giving 4,044 consecutive-day pairs. Day-to-day overlap of cited sources came in at Jaccard 0.34 to 0.42 with RBO 0.21 to 0.26, meaning roughly two thirds of cited sources change overnight. Brand-level overlap was higher at Jaccard 0.45 to 0.59. Repeated queries fired within a single 24-hour window produced source overlap of 0.32 to 0.43, close to the day-to-day figure, which means most of the apparent day-over-day churn is within-day randomness rather than anything changing on the web.
That study's sampling recommendations are concrete: 7 runs per prompt per day to get standard error below 0.10 on per-brand detection, 8 runs for source coverage, 10 days for the window to reach the same threshold and 24 days to halve it. The 2026 critical survey of GEO research repeats the same guidance, citing daily source-level Jaccard of 0.34 to 0.42 and recommending "seven to eight repetitions per prompt as a starting point".
The two studies do not disagree with each other. They measure different units. Set overlap asks which sources appeared today versus yesterday, and sources churn heavily. Per-cell rates ask whether a given brand shows up for a given prompt across the whole period, and that is close to binary. Sources are a sample from a large pool. Your presence in an answer usually is not.
#Three metrics, three cadences
Report mention as a state with a change log rather than a trend line, and name the prompts whose state changed. Treat the cells in the volatile band as the working set for content work, since they are the ones where an intervention can move anything.
Hold sentiment to a wider threshold and a longer window than mention, and stop publishing week-over-week tone deltas on small prompt sets. When sentiment does move past that threshold, read the answer text before accepting the label. The survey's recommendation is that audits pair citation rates with "a matrix covering tone, attribution accuracy, and factual support", which puts tone beside verification rather than alone on a chart.
Watch the denominator underneath all of it. The same survey notes that 57.8% of ChatGPT repetitions did not activate web search in one dataset, and that discarding responses without search or citations creates selection bias. Sentiment computed only over answers that mentioned you, over runs that happened to search, is three filters deep before anyone looks at it.
#What these numbers do not support
The tier figures rest on small and uneven groups, 11, 36 and 55 brands respectively, hand-classified without inter-rater validation, drawn from a convenience cohort skewed toward SaaS, retail, fintech and Indian direct-to-consumer brands, with one model variant per engine and no randomized comparison. The sampling study ran from Swiss IP addresses with German-language prompts and detected brands by substring matching, so synonyms and abbreviations were missed.
Use the ratio, not the levels. Whether sentiment is 6.7 times noisier than mention in your category is unknown. That it is far noisier, and therefore needs more samples and wider thresholds before anyone reacts to it, is the part that travels.
Related field notes
September 24, 2026 · 6 min
Google pays for grounding, not for links
Google's AI contribution pilot pays when a page shapes an answer, not when it is linked afterward. That rule says a citation count measures the wrong thing.
September 23, 2026 · 5 min
A browser agent is not a crawler
Agentic browsing runs inside the user's own session, so robots.txt, bot allowlists and crawler analytics all miss it entirely.
September 22, 2026 · 4 min
An MCP endpoint is not a discovery channel
NLWeb and MCP make your site answerable by agents that already found you. Nothing on the open web is hunting for a /mcp route yet.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.