Licensing deals don't buy AI visibility
Crawlmind Engineering··5 min read
An AI content licensing deal is a commercial agreement letting an AI company use a publisher's archive, and the best public evidence that such a deal buys extra citations is a correlation with a plausible non-causal explanation.
The number in circulation comes from a joint study by Press Ranger and Otterly.ai, released in August 2026. They matched a month of citation data against every confirmed publisher agreement they could find and reported that pages from OpenAI-licensed publishers earned 48% more citations on ChatGPT than pages from unlicensed ones, in their joint announcement. That figure has since been repeated as though a deal were a distribution channel you can buy. The study's own numbers argue otherwise.
#What the 48% is measured over
The study captured 129.3 million citations in June 2026 across seven platforms, and mapped 314 publisher domains to 91 confirmed agreements, according to Otterly's writeup. On ChatGPT, licensed publishers averaged 10.2 citations per cited page against 6.9 for unlicensed publishers. Across all seven platforms the split was 10.7 versus 7.3, a 46% gap, per the same writeup.
The unit is citations per cited page. That normalization is there for a good reason: it stops the Times of London from beating a trade blog purely on volume. It also does something less obvious. Conditioning on pages that were already cited at least once removes the entire question a GEO team actually cares about, which is whether a page enters the candidate set at all. A publisher whose citations concentrate on a few evergreen pages scores high on this metric. A publisher with the same total citations spread thinly across a large archive scores low. Neither pattern tells you the licensing deal moved anything.
The exclusivity gradient looks like the strongest evidence for causation. Publishers signed only with OpenAI took 112% more citations per page on ChatGPT than unlicensed publishers, per Otterly. Exclusivity also correlates with the kind of publisher OpenAI approached first, and the study is explicit that it reports association rather than causation.
#The platform contrast is the tell
If a licensing deal caused citations, the effect should show up on the platform doing the licensing, for each licensor. It does not.
Google-licensed publishers were cited slightly less often in AI Overviews than unlicensed ones, and Perplexity licensees sat at roughly parity, as Elmo's analysis of the study points out. Microsoft Copilot showed the widest gap of any platform at 85%, though Otterly does not break that down by which company held the agreement.
A simpler variable fits all four results. "Licensed publisher" is largely a proxy for "prominent, English-language news brand." Those brands do well on platforms whose retrieval leans on major news domains and worse on platforms that pull from a wider source mix. Licensing status is what you can observe; prominence is what is probably doing the work. The study reports that the top five media groups control 69% of licensed citations, 88.7% of which come from English-language publications and 73.0% from North American ones, all in the Otterly data. That is a description of a small, already-dominant set of domains.
There is also a timing problem. Any single month is a thin slice of a source layer that turns over quickly, and this was a June 2026 snapshot released as a vendor press release rather than a peer-reviewed paper, a limitation Elmo states plainly. The Press Ranger and Otterly.ai announcement is the primary artifact; there is no published methodology document to check the publisher-matching criteria against.
#What actually gets cited from licensed publishers
The most useful finding in the study has nothing to do with deals. Of the citations going to licensed publishers, 58.8% pointed at commercial evergreen content and 38.1% at best-of lists, while dated news articles took 5.1% and wire copy under 1%, per Otterly.
News publishers are being cited for their buying guides, not their journalism. That layer, the "best X for Y" roundup, is a layer any brand competes in already. It is also the layer where inclusion is an editorial decision made by a human, not an algorithmic one made by a retriever.
News as a whole is a minority of the citation pool. Media content took 7.2% of the 129.3 million citations measured, ranging from 14.0% on Claude to 8.3% on ChatGPT, according to Otterly. Most of what AI engines cite is not news at all, which caps how much any publisher-licensing dynamic can explain about your own visibility.
#The reachable pool is unlicensed
Trade and niche media collected 213% more AI citations than mainstream media across 16 industries in the same dataset, winning in 15 of the 16, per Otterly. The licensing overlap runs the other way there: 54.7% of citations within mainstream media went to licensed publishers, against 17.2% within trade media.
That is the practical read. The part of the citation pool most likely to mention a specific B2B product is trade press, and roughly five sixths of it has no AI licensing agreement at all. Earned coverage in a vertical publication is a lever you can pull. A licensing deal is not.
#Being cited is not being represented correctly
One more result belongs next to the 48%. The Tow Center at Columbia tested 200 traceable quotes from 20 publishers against ChatGPT Search and found 153 responses that were wrong or partly wrong about the source, with the system acknowledging it could not answer only seven times, in the CJR report. Publishers with licensing agreements and open crawlers were among those misquoted; the New York Times, which blocked OpenAI's crawlers, still had quotes attributed to it, sometimes via plagiarized copies on other sites.
Separately, Tollbit's State of the Bots reporting found OpenAI-licensed publishers seeing a ChatGPT clickthrough rate close to seven times that of unlicensed publishers, as covered by Press Gazette. Whatever a deal does or does not buy in citation count, citation count and downstream traffic are separate measurements, and so is whether the answer described you accurately.
#What to do with this
Treat the 48% figure as a finding about publishers, not a strategy. Three things follow for a brand tracking its own AI visibility.
Measure entry into the candidate set separately from citation intensity. Share of prompts where you appear at all answers a different question than citations per cited page, and only the first one tells you whether you are in the running.
Weight trade and vertical coverage above mainstream placement in your PR planning, because that is where the citations concentrate and where licensing status does not sort the field.
Verify what the answer says, not only that a link appeared. A citation that misattributes your claim is a measurement event, not a win.
Related field notes
September 24, 2026 · 6 min
Google pays for grounding, not for links
Google's AI contribution pilot pays when a page shapes an answer, not when it is linked afterward. That rule says a citation count measures the wrong thing.
September 23, 2026 · 5 min
A browser agent is not a crawler
Agentic browsing runs inside the user's own session, so robots.txt, bot allowlists and crawler analytics all miss it entirely.
September 22, 2026 · 4 min
An MCP endpoint is not a discovery channel
NLWeb and MCP make your site answerable by agents that already found you. Nothing on the open web is hunting for a /mcp route yet.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.