AI engines cite synthetic pages from the tail
Crawlmind Engineering··5 min read
A synthetic source is a page cited in an AI answer whose text was mostly written by a language model, and three independent studies now find that these pages make up a real share of what answer engines cite, concentrated in the long tail rather than on well-known domains.
That matters for anyone doing GEO. The pages you compete with for a citation are not only your known rivals and the big publishers. On many questions, especially narrow ones, the page that wins the citation is a generated article on a site nobody on your team has heard of.
#What the Northwestern audit found
The most direct evidence comes from Allaham and Diakopoulos at Northwestern, submitted to arXiv on May 22, 2026. They ran 712 English-language queries on politics, health and environment through ChatGPT, Copilot, Gemini and Perplexity, collected 26,266 unique cited URLs across 7,675 domains, and scraped 19,154 of them. They then classified each page with Pangram, an AI-text detector they picked for its low false positive rate.
About 16% of the scraped cited sources (3,056 pages) were classified as likely or highly likely AI-generated. The split by engine was wide:
| Engine | Cited sources classified as AI-generated |
|---|---|
| Copilot | 27.8% |
| Gemini | 14.7% |
| Perplexity | 9.4% |
| ChatGPT | 7.3% |
The distribution finding is the one to remember. The top 25 cited domains accounted for only 2.9% of the AI-generated sources, and the other 97.1% sat in the long tail. Across all citations, 59.1% of domains were cited exactly once. Synthetic pages do not arrive through a few big offenders. They arrive as a large number of small sites that each win a citation or two.
The authors call their numbers a lower bound. On a set of known AI content farm domains, Pangram missed 31.4% of the generated text. The audit also covers only English, US-relevant, single-turn queries, and 27.1% of cited URLs could not be scraped at all (PDFs, images and blocked pages).
#Two other studies point the same way
Originality.ai analysed AI Overview citations for 29,000 YMYL queries (health, finance, legal and politics) and published the results on October 28, 2025. With its own detector, it classified 10.4% of cited documents as AI-generated, 74.4% as human-written and 15.2% as unclassifiable. The split by ranking position is telling. Among citations that also appeared in the top 100 organic results, 7.7% were classified as AI-generated. Among citations from outside the top 100, the figure was 12.8%. Just over half of all citations (52%) came from outside the top 100.
Graphite's October 2025 study looked at 31,493 keywords across 10 categories using Surfer's detector. It found 14% of articles in Google results were AI-generated, but only 7% of top-ranked articles. For ChatGPT and Perplexity citations, the AI-generated share was 18% on both.
The three studies use different detectors, different query sets and different engines, so the percentages should not be compared directly. The direction is consistent, though. The share of synthetic pages is lowest at the top of organic rankings and higher in the pages AI engines pull from further down or outside them.
#Why the tail is where this happens
An answer engine needs a supporting page for every claim in its answer. On a broad query, strong candidates are plentiful and the engine can pick well-known sources. On a narrow question (a specific dosage interaction, a regional regulation, a niche product comparison), there may be few human-written pages that address the exact phrasing. A generated article that matches the question closely can be the best available fit for retrieval, even on a domain with no history.
Classic ranking has spent two decades building signals that push such pages down: links, site reputation, spam classifiers. Google names the pattern in its spam policies, which define scaled content abuse as "many pages are generated for the primary purpose of manipulating search rankings", and give "using generative AI tools" to produce pages "without adding value for users" as an example. Retrieval for answer grounding appears to reach past the part of the index where those signals apply most strongly. The Originality.ai split is consistent with that reading: more of the out-of-top-100 citations were classified as generated than the in-top-100 ones.
#What to do with this
Map the tail you actually lose. If your citation tracking only covers head terms, you will mostly see recognisable competitors. Run the narrow, specific questions your buyers ask and record which domains get cited. Where the winner is an unfamiliar site with a generated page, that query has no strong human source, and it is open for you to take.
Answer the specific question on a page you would stand behind. The advantage a generated page has on a tail query is fit, not quality. Match the fit with a page that states the answer plainly near the top, uses the terms the question uses, and adds something a model cannot invent: your own data, a worked example, a named expert, a date. That combination is harder for a generic page to match.
Do not respond by generating pages at the same scale. A tactic that works only because retrieval filters are weaker than ranking filters depends on that gap staying open, and Google's scaled content abuse policy already applies to the organic side. Graphite also notes it did not evaluate heavily edited AI-assisted content, so the evidence does not tell you that model-drafted pages with real human editing underperform. It does tell you that unedited volume is the pattern these studies are counting.
Classify cited sources before you benchmark against them. When you build a list of "who gets cited instead of us", tag each domain as publisher, competitor, community, reference, or unknown and unverified. Copying the structure of a generated page because it wins a citation today puts you in the category engines are most likely to filter next.
Treat detector labels as estimates. Every study here depends on an AI-text classifier, and the Northwestern team measured a 31.4% miss rate on known content farms. A single page labelled as AI-written is a signal to read it, not a verdict.
#The takeaway
Engines differ a lot here. In the Northwestern audit, Copilot's synthetic share was nearly four times ChatGPT's, so one blended "AI visibility" number hides very different competitive sets. The practical reading across all three studies is the same, though. Synthetic pages win citations where nothing better exists. The long tail of specific questions is where that gap is widest, and it is where a well-sourced human page has the clearest chance to replace them.
Related field notes
September 25, 2026 · 5 min
Your AI citations have a shelf life
New studies track AI citations over weeks and months. Most cited pages get replaced, and the rate depends heavily on the engine and the URL.
September 25, 2026 · 5 min
Engines are learning to discount GEO rewrites
Two September 2026 papers build filters against manipulative GEO. They catch attacks by style, so honest pages that copy the style pay a small tax.
September 25, 2026 · 5 min
Your rank-to-citation chart is mostly selection
A September 2026 causal audit found top-ranked sources cited far more often, yet moving a page up barely changed its odds. Relevance did the work.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.