Your rank-to-citation chart is mostly selection
Crawlmind Engineering··5 min read
Citation allocation is the step where an AI answer engine, holding several retrieved sources that support the same claim, decides which of them to cite and how often.
Most GEO reporting treats that step as a function of position. Pages that retrieve first get cited more, so the advice follows: get to the top of the retrieval list and the citations will come. A causal audit posted to arXiv on September 14 tested that advice directly and found the observed gradient is mostly something else.
#The chart everyone draws
The paper is CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search by Sriram Selvam and Anneswa Ghosh. The setup is realistic. A GPT-5.4 agent answered 129 everyday queries, issuing 346 search calls over 1,490 unique documents, using Exa as the retrieval provider with five results per call. Answers carried a median of six distinct cited documents.
In those authentic transcripts, position looks decisive. Documents first shown at rank 1 were cited 85.1% of the time, against 42.8% at rank 5, a gap of 42.3 points. That is the chart most citation dashboards produce, and on its face it says position is worth a great deal.
#What happened when they moved the page
The authors then held everything fixed and swapped the order of evidence-matched documents: two pages that support the same fact, with only their positions exchanged. In the scaled replay, the position effect shrank to 7.9 points. In 56 held-out pairs swapped in isolation, it was exactly 0.0 points.
The authors call this a deflationary result: the descriptive gap is about five times the controlled estimate and did not survive held-out confirmation.
The reason is ordinary selection. A search provider puts more relevant, higher-quality pages first, and the agent chose the queries that produced those rankings. The rank 1 page is cited more because of what it is, and it sits at rank 1 for the same reason. Correlating rank with citation mixes the two, and the chart credits position for work that relevance did.
None of this says position never matters to a language model. The 2023 paper Lost in the Middle showed models use information at the start or end of a long context better than information buried in the middle. The CITECHOICE finding is narrower: with five short results per call, swapping two matched sources moved citations far less than the observational data implied.
#The same confusion at the SERP level
The rank story has a second version for Google. In July 2025, Ahrefs found about three quarters of AI Overview citations came from the top 10 organic results. Its March 2026 follow-up across 863,000 keyword SERPs and 4 million AI Overview URLs put that share at 37.9%, with 31.2% from positions 11 to 100 and 31.0% from beyond the top 100.
That is not a causal study either, but it points the same way. Organic rank and AI citation are both downstream of relevance to the question the system actually ran, and Google does not expose the fan-out queries it generates. Treating page one as the lever misreads a correlation that is weakening on its own.
#Structure moved credit, not admission
The paper also tested presentation. The authors generated a prose rendering and a structured rendering (headings, short paragraphs, lists or a table) of the same source text, required to keep the same claims, quantities, caveats and attribution within 25% of each other in word count.
Structure raised the number of citation markers a page received by 0.50 per answer, with a 95% interval of 0.20 to 0.84. It did not clearly change whether the page was cited at all: that estimate was 4.5 points with an interval from minus 1.4 to plus 10.4, which the authors report as unresolved rather than zero.
Two caveats from the paper matter. The two versions were not word-identical, so this measures a rewrite package, not list markers alone. The authors also say no stable list-marker or rank recipe emerged, and describe the result as an attribution-sensitivity warning, not an optimization tactic. That fits an earlier finding covered here, that formatting-only edits barely move citation odds. Structure seems to help a page that is already in the answer collect more of the credit. It does not get a page into the answer.
#One answer is too noisy to read
The most practical number in the paper is the noise floor. Regenerating the same transcript, binary citation decisions agreed 85.0% of the time, so roughly one decision in seven flipped with no change to any input. Exact citation counts matched only 61.7% of the time, and decoding noise accounted for an estimated 45% of the variance in a single-draw effect.
If you compare one answer before a change with one answer after it, most of what you see could be that noise. The same point shows up in how many reruns a prompt needs, which we covered in the fifth rerun buys almost nothing.
#What to do with this
- Stop reading rank-to-citation charts as a lever. Treat them as a description of which pages are relevant, not proof that moving up will earn citations. If you want a causal answer, change one page and hold a matched control page constant.
- Compete on the fact, not the slot. Across the pairs where two pages supported the same fact, swapping order did little. The page that holds a specific, checkable fact others lack is the one that competes for admission.
- Track admission and credit separately. Whether a page is cited and how many times it is cited responded differently in this study. A report that merges them hides which one changed.
- Repeat before you conclude. Run each prompt several times per condition and report agreement, not a single answer.
- Check what your content becomes after extraction. The authors note that extractors keep headings, lists and table rows in the text the model sees. How your page survives that step is part of how it gets credited.
The study has limits the authors state: one retrieval provider, one answering model, five results per call and bounded page text. Other engines may weigh position more. The method is the useful part. Before crediting position for a citation, check whether the page would have been cited from any slot.
Related field notes
September 25, 2026 · 5 min
Your AI citations have a shelf life
New studies track AI citations over weeks and months. Most cited pages get replaced, and the rate depends heavily on the engine and the URL.
September 25, 2026 · 5 min
Engines are learning to discount GEO rewrites
Two September 2026 papers build filters against manipulative GEO. They catch attacks by style, so honest pages that copy the style pay a small tax.
September 24, 2026 · 6 min
Google pays for grounding, not for links
Google's AI contribution pilot pays when a page shapes an answer, not when it is linked afterward. That rule says a citation count measures the wrong thing.
Share or discuss
New posts, no spam. Roughly monthly. Unsubscribe with one click.