Your paywall hides the body, not the record
Crawlmind Engineering··5 min read
A paywall is an access rule applied at fetch time: it changes what an AI engine can read about you, and it does not by itself decide whether your URL can appear in an answer. Those are two different layers, they are measured by different studies, and that is why the published evidence on paywalls and AI visibility looks like it contradicts itself.
It mostly does not contradict itself. It measures different things and reports both as "citations."
#The strongest version of the paywall claim
The headline result comes from 5W's "Paywall Penalty" study. Across a 40-query test, it reports hard-paywall and metered publishers at 0% of AI-retrieval citations, freemium at 8.7% and open-web publishers at 91.3%. Forbes alone accounted for 58% of that freemium share. The publishers named at zero were the Wall Street Journal, Financial Times, Bloomberg, the New York Times, the Washington Post, the Economist and the Atlantic.
Read the instrument before you read the result. The 40 queries covered six categories, from hard-news retrieval to best-X recommendation, and the whole test ran on a single retrieval instrument: Claude with real-time web search. Cohort assignments were fixed in advance from each publisher's access posture. A second wave at 500 queries across five engines is described as still under development.
Forty queries on one engine is a query mix, not a structural law. A zero at that sample size means "did not surface in these forty," which is a weaker statement than "cannot be cited."
#The same vendor's larger dataset points the other way
5W also publishes a consolidated citation index built from more than 680 million individual citations across six published studies between August 2024 and April 2026. In that index, Claude is described as favoring the New York Times, the Atlantic, the New Yorker and the Economist, with 36% of its journalism citations from the past 12 months versus 56% for ChatGPT.
Same vendor, same engine, opposite direction. Four of the seven publishers reported at zero in the 40-query test are named as the engine's preferred sources in the larger index.
That is not a scandal. It is a denominator. Hard-news retrieval and best-X recommendation prompts pull from different publisher pools, and a prestige-editorial preference shows up on the prompts where prestige editorial is the relevant corpus.
The per-engine picture in the same index is graded rather than binary. On ChatGPT, the published ranking puts the New York Times at #14 and Bloomberg at #15 in the top-50 source list, with the Wall Street Journal at #21 and the Financial Times outside the top 50. Those are modest positions for newsrooms of that size, and they are not zero.
Low share is the defensible claim. Structural zero is not. A publisher that ranks fourteenth on one engine and scores zero-for-forty on another in the same year is being measured, not excluded.
#Blocking is the other lever, and it barely moves citations
If access controlled citation, an outright crawler block would be the cleanest test, because a block is unambiguous in a way a metered paywall never is.
BuzzStream, working with Citation Labs data, analyzed 4 million citations from 3,600 prompts across ChatGPT, Gemini, AI Overviews and AI Mode in ten industries, published April 2026. Among the top 50 news sites, it reports 88.2% blocking GPTBot and 92.3% blocking Google-Extended, with 70.6% blocking ChatGPT-User and 82.4% blocking OAI-SearchBot. Those blocks did not clear the sites out of the answers. In the same dataset, 95.4% of the GPTBot-related citations came from sites blocking GPTBot, CNBC appeared 1,298 times while blocking broadly, and Yahoo appeared in nearly 30,000 citations while blocking Google-Extended.
The mechanism is mundane. A citation can be minted from a search index record the engine already holds, from a headline and a result snippet, or from another site's summary of your article. None of those require reading your body text on the day of the query. That is also why a citation is not evidence that the engine understood you, a gap we have written about in cited but not absorbed.
Treat the BuzzStream figures as directional rather than exact. The study does not publish a strict definition of what counted as a citation or where in the response it had to appear, and a looser definition inflates retention.
#The cost of a block shows up in traffic, and that number was revised
The clearest measured effect of blocking is on visits, not citations. Zhao and Berman, at Rutgers Business School and Wharton, studied 30 major newspaper publishers, expanded to the 500 largest, over November 2022 to May 2024. The current version reports a 7% decline in total traffic within six weeks of blocking, concentrated among the top 50 publishers and weaker for smaller ones. They also find around 75% of top publishers blocking LLM crawlers from mid-2023, a 31.2% fall in article volume, and increases of 68.1% in interactive elements and 50.1% in advertising technologies.
The figure in wide circulation is 23%, not 7%. That came from the earlier draft, measured monthly; the current version measures weekly visits and lands near 7%. If you are citing this study to justify a policy, cite the version you actually read.
#Declaring the gate is not the same as closing it
There is a separate confusion worth clearing. Google's paywalled-content structured data, using isAccessibleForFree set to false with hasPart and a cssSelector, exists so that gated sections can be differentiated from cloaking, which is a spam policy violation. The documentation is explicit that it "only applies to content that you want crawled and indexed".
So the markup is a declaration about a page you are still inviting Google to read. It does not open the gate for anyone, it does not close it, and it has nothing to say to engines outside Google's crawl. Adding it is correct hygiene for a paywalled publisher. It is not a visibility lever.
#Measure the three layers separately
The reason these studies appear to fight is that "AI visibility" is being used for three questions with three different answers.
Fetch is the first: does the bot get a 200 and a body? That lives in your server logs, split by user agent, and it is the only one of the three you control directly.
Citation is the second: does your URL appear in the answer? That is measurable with citation tracking and, on the evidence above, it survives both paywalls and blocks far better than most publishers expect.
Absorption is the third: does the answer carry your claim, or a competitor's version of it? That is the one a paywall actually threatens, because an engine that can only read your headline will reach for whoever wrote the accessible explainer. Counting appearances without checking what was said is the failure mode we covered in count verified mentions, not mentions.
The decision in front of a gated publisher is not paywall or no paywall. It is which of those three layers you are prepared to lose, and whether your reporting can currently tell them apart.
Related field notes
September 24, 2026 · 6 min
Google pays for grounding, not for links
Google's AI contribution pilot pays when a page shapes an answer, not when it is linked afterward. That rule says a citation count measures the wrong thing.
September 23, 2026 · 5 min
A browser agent is not a crawler
Agentic browsing runs inside the user's own session, so robots.txt, bot allowlists and crawler analytics all miss it entirely.
September 22, 2026 · 4 min
An MCP endpoint is not a discovery channel
NLWeb and MCP make your site answerable by agents that already found you. Nothing on the open web is hunting for a /mcp route yet.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.