We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

Your AI Overview loss depends on the baseline

Crawlmind Engineering··5 min read

An AI Overview traffic loss estimate is the gap between the search traffic a page received after AI Overviews launched and the traffic it would have received without them, and that second number is never observed, so it has to be borrowed from a comparison group.

The choice of comparison group is where most of the answer comes from. One research paper shows this more clearly than any vendor dashboard, because it published the same question with different baselines and got answers three times apart.

#One paper, two headline numbers

Mehrzad Khosravi and Hema Yoganarasimhan at the University of Washington used Wikipedia as a natural experiment. Wikipedia publishes the same article in many languages, and Google rolled AI Overviews out to English searches before most other markets. If English articles lost search traffic relative to their own German or French versions, the gap is a reasonable estimate of what AI Overviews did.

An earlier version of the paper compared 52,262 English articles against their Hindi, Indonesian, Japanese and Portuguese versions and reported an English decline of roughly 15%. The same version broke that down by topic: culture articles fell 19.6%, geography 16.6%, history and society 9.9%, and STEM 7.4%.

The current version, posted September 2, 2026, changed the design. It now uses monthly external-search referrals as the outcome and German and French editions as controls, since those markets did not get default AI Overviews until early 2025. The headline estimates are a 5.45% decline against German and 4.82% against French, built on roughly 500,000 matched article pairs per comparison.

The same current version still reports an English versus Japanese comparison through July 2024 that yields a 16.53% decline. The authors present it as confirming the direction of the effect. It also shows how far the magnitude moves when only the control group changes.

None of this makes the paper unreliable. It makes it honest. The authors ran placebo tests, parallel-trend checks and sensitivity bounds, and the direction held in every specification. The size did not.

#Why the baseline moves the answer

A difference-in-differences estimate assumes the control group would have moved exactly like the treated group if nothing had happened. Every choice below breaks that assumption a little differently.

The control market has its own trends. Japanese and Indonesian search behavior in 2024 was not a copy of American behavior minus AI. Mobile share, competing apps and local search engines all drift on their own schedule. German and French are culturally closer to English Wikipedia usage, which is one reason the later estimate is smaller.

The outcome metric changes what counts. Total pageviews include visits from bookmarks, internal links and other sites. External-search referrals isolate the channel AI Overviews actually touch. The paper's own methods note says it avoids using other referral sources as a control because AI Overviews may affect traffic to those websites too.

The model specification changes the scale. In the earlier version, the same data produced a 3.47% change under a Poisson specification and an 8.1% decline under a weighted log model, against the roughly 15% headline. Log models on traffic data are sensitive to the many small pages with near-zero counts.

The bot filter changes the history. In October 2025 the Wikimedia Foundation reported that human pageviews were down roughly 8% against the same months in 2024, but only after it found that unusually high traffic in May and June 2025, mostly from Brazil, came from bots built to evade detection, and reclassified March through August. The Foundation itself warned the revised series has to be read with care because its detection rules differ across time periods. Any before-and-after comparison that straddles a bot-filter change is partly measuring the filter.

#The same feature can raise traffic elsewhere

A baseline problem also shows up across site types. A separate study of 105,012 subreddits by Zhang, Cui and Zhang compared communities Google surfaces in AI Overviews against adult communities it excludes. After the August 2024 international expansion, daily comments in the exposed communities rose 12.0% and unique commenters rose 12.4%, with the largest proportional gains in small and medium communities. The same paper found that AI Mode, launched in August 2025, largely erased the extra gain for experience-based communities.

So "AI Overviews cut traffic" is true for Wikipedia and false as a general rule for Reddit participation. A single industry-wide number does not transfer to your site.

#What a site owner can take from this

The randomized experiments we covered in AI click loss now has a causal number put the click loss at 39.8% on informational queries where the overview sits at position zero, with no measurable effect on navigational or transactional queries. That is a per-query number. The Wikipedia estimates are per-site averages over every query mix. Both can be right at once, which is exactly why you should build your own baseline instead of borrowing one.

A workable approach for a single site:

  1. Pick a control you own. Pages that rarely trigger an AI Overview (branded, navigational, transactional) make a better baseline than last year's traffic, because they share your seasonality, your domain authority and your technical changes.
  2. Measure the channel, not the site. Use organic search clicks from Search Console, split by query type, rather than total sessions from analytics. Direct, referral and AI-assistant traffic have their own trends.
  3. Freeze your bot filter. If your analytics vendor, CDN or WAF changed its bot classification during the window, restate the earlier period under the new rules or exclude it. Otherwise the change in filtering reads as a change in demand.
  4. Report a range. Run the comparison against at least two controls and publish both numbers. If they differ by a factor of three, as they did for Wikipedia, that spread is the honest answer.
  5. Segment by page intent. Informational pages and product or pricing pages behave differently. An average across both hides the one that is actually losing clicks.

Crawlmind's AI visibility reports track citation presence per query and per engine for this reason: whether a page is cited in the answer is the variable that differs between treated and untreated queries, and it is the one a site can act on.

The Wikipedia paper is useful mainly as a method. Its authors changed the baseline and the number moved by a factor of three, and they published both. If your own AI Overview loss estimate comes with only one number and no stated control, you don't know what it is measuring yet.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.