We respect your privacy.

We use strictly necessary cookies to keep you signed in and to protect against CSRF. With your permission we also use a small amount of first-party analytics to improve the product. We do not sell your data and we do not use third-party advertising trackers. See our cookie policy and privacy policy .

← All posts

What an AI licensing deal actually buys

Crawlmind Engineering··5 min read

An AI content licensing deal is a paid agreement letting an AI company use a publisher's archive, and the evidence that it also buys citation placement is considerably weaker than the headline number going around suggests.

The headline number is real. On 20 August 2026, Press Ranger and OtterlyAI released an analysis of 129.3 million citations collected across seven AI platforms during June 2026, matched against every confirmed licensing agreement between AI companies and news publishers. On ChatGPT, pages from OpenAI-licensed publishers averaged 10.2 citations each against 6.9 for pages from unlicensed publishers, a 48% premium. Publishers that signed with OpenAI and nobody else did better still, at 112% more citations per page.

Read quickly, that says money buys visibility. Read the rest of the same dataset and it says something narrower.

#The asymmetry that breaks the causal reading

If a licensing deal worked as a placement lever, the effect should show up on the platform belonging to the licensor. Across the seven platforms measured, it shows up on exactly one.

Google's licensees were cited slightly below comparable unlicensed publishers on Google AI Overviews. Perplexity's licensees landed at parity with unlicensed publishers on Perplexity. Only the OpenAI cohort shows a home-platform premium, and that cohort also carried 46% more citations per cited page across all seven platforms combined, including platforms it has no agreement with.

That last detail is the one to sit with. A publisher's OpenAI contract cannot plausibly be raising its citation rate on Gemini or Claude. So whatever is producing the cross-platform lift is a property of the publishers themselves, not of the contracts. Once you accept that for six platforms, the burden of proof for the seventh gets much heavier.

Three explanations fit the pattern. OpenAI may handle licensed content differently while its competitors do not. Or OpenAI signed a different class of publisher than Google and Perplexity did. Or ChatGPT's retrieval favors exactly the kind of large, English-language news brand that tends to end up on a deal list. The published analysis cannot separate them, and it does not claim to. The write-ups are explicit that the study does not show licensing deals caused the differences.

#Citations per cited page is a conditional metric

The measured unit deserves a second look, because it is not the thing most people think they are reading.

Citations per cited page counts how many times a page gets cited given that it was cited at all. Pages that never surfaced contribute nothing to either side of the comparison. So the metric describes citation density among content that already cleared the retrieval bar. It says little about the odds of clearing that bar in the first place.

This matters for anyone converting the finding into a strategy. A density premium among pages that are already visible and a lift in the odds of becoming visible are different claims with different implications, and only the first one was measured. It is the same denominator trap that makes published AI-citation studies appear to contradict each other when they are actually reporting different quantities.

#Who got signed, and when

Selection sits underneath all of it. The licensed cohort is roughly 40 OpenAI partnerships, and that population is not a random draw. Press Gazette's read is that current AI licensing skews heavily toward English-language premium journalism from large media businesses, with small, non-English and independent outlets largely shut out. Those are the same publishers a retrieval system would tend to favor on brand, domain authority, and archive depth with no contract in place.

The comparison group is also unspecified. The analysis has not been peer reviewed, has no published methodology, and does not explain how the unlicensed publishers were chosen. It is a press release from two companies that sell AI-visibility products, and it should be weighed as one.

None of that makes it worthless. It makes it a correlation with an obvious confound and a single-month window, which is a reasonable thing to publish and an unreasonable thing to build a budget on.

#What the data does support

Strip out the causal claim and a durable finding remains: OpenAI-licensed publishers now draw 57.9% of all their AI citations from ChatGPT alone, while unlicensed publishers show a more even spread across engines.

That is a concentration result, and concentration is a risk as much as a win. A publisher taking more than half its AI visibility from one engine is exposed to that engine's next retrieval change in a way a diversified publisher is not. There is a real revenue case for the deals, separately: Tollbit's Q2 2025 State of the Bots report found licensed publishers achieved a ChatGPT clickthrough rate almost seven times higher than unlicensed ones. Clickthrough behavior on cited links is a different mechanism from citation selection, and it is worth keeping the two apart rather than folding both into one story about paid placement.

#The disclosure gap this exposes

The speed with which a correlational press release became "preferential treatment confirmed" across marketing blogs is itself the interesting artifact. Nobody outside these companies can currently check the claim, in either direction.

An ICML 2026 Position Track paper, Generative Engine Optimization Creates Underexamined Risks, names this directly. It identifies undisclosed commercial influence embedded in evidence and reasoning as one of three governance risks in answer engines, alongside concentrated influence arising from low contestability, and calls for high-precision disclosure plus black-box auditing of material influence. Low contestability is the operative phrase. When a ranked list of links buries you, you can see it happen. When a synthesized answer omits you, there is no ranking to inspect and no appeal.

Until disclosure of that kind exists, questions like "do licensing deals affect citation selection" stay unanswerable from outside, and outside is where everyone reading this sits.

#If you cannot sign a deal

Most organizations cannot. The practical takeaways are unglamorous.

Do not treat licensing as a GEO lever you are missing out on. The finding is about news publishers, on a news-heavy citation corpus, using a metric conditional on already being cited. It does not generalize to a B2B product page.

Do track your citation distribution across engines rather than a single blended score. If one engine supplies most of your visibility, you have the licensed publishers' concentration profile without the licensing revenue that compensates for it.

And treat any single-month, single-vendor citation study, including the ones vendors like us could run, as a direction rather than a magnitude. The direction here is that large incumbent news brands are consolidating AI citation share. That was already true before anyone signed anything.

Related field notes

Share or discuss

Field notes in your inbox

New posts, no spam. Roughly monthly. Unsubscribe with one click.