[The Engines]
How Do AI Search Agents Decide Which Pages to Cite?
A causal audit of a GPT-5.4 search agent found that the citation gap between the top-ranked page and the fifth-ranked page shrank from 42.3 points to near zero once researchers actually tested cause and effect.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
A new causal study of a GPT-5.4 search agent found that top-ranked pages looked far more likely to get cited than fifth-ranked pages, an 85.1% to 42.8% gap in the raw data. When researchers actually swapped page rankings to test cause and effect, the real position effect dropped to 7.9 points, then to statistically zero in a follow-up test. Rank alone is not the lever it appears to be; what a page says, and how it is structured, carries more of the weight.
What did the new study actually find?
Researchers Sriram Selvam and Anneswa Ghosh published a preprint on arXiv on September 14, 2026, titled 'CiteChoice: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search.' They tested a GPT-5.4 search agent running on the Exa search API, pulling 129 real query transcripts and isolating 113 document pairs where two separate pages supported the identical fact. A blinded review confirmed 103 of those pairs.
Instead of just watching which pages got cited, they intervened. They swapped which page ranked first in the result set and swapped which page used structured formatting versus plain prose, then measured what actually changed in the agent's citation behavior. That distinction, between what correlates with citation and what causes it, is the whole story.
What changed between the raw data and the controlled test?
In the unmanipulated dataset, pages ranked first got cited 85.1% of the time. Pages ranked fifth got cited 42.8% of the time. That is a 42.3 percentage point gap, and on its face it confirms the assumption most SEO and GEO teams already operate on: rank high, get cited.
Then the researchers ran the actual experiment. They held content constant and swapped positions to isolate cause from correlation. The real, controlled effect of position dropped to 7.9 percentage points, about a fifth of what the raw numbers implied. A second replication test, using 56 pairs, found no measurable position effect at all: a 95% confidence interval spanning -5.4 to +5.4 points, statistically indistinguishable from zero.
| Measurement | What it shows | Effect size |
|---|---|---|
| Raw, unmanipulated data | Rank 1 pages vs. rank 5 pages, observed citation rate | 42.3 percentage point gap (85.1% vs. 42.8%) |
| Controlled reordering test | Same content, positions swapped to isolate causation | 7.9 percentage point gap |
| Replication test (56 pairs) | Second controlled run | 0 points, statistically indistinguishable from zero (95% CI: -5.4 to +5.4) |
Why did the raw numbers lie?
Pages that already rank first tend to come from domains with stronger topical authority, tighter query match, or better editorial signal. Those same qualities make a page more citable on their own. Position and citation both trail the same underlying causes, quality and relevance, without position itself doing the causing. Strip that overlap away, as the controlled test did, and most of the apparent 'rank effect' disappears.
This is the trap raw dashboards set for anyone tracking AI visibility. A tool that shows citation rate by SERP position will always show a gap. Whether that gap reflects a lever you can pull, or just two effects sharing a cause, requires a test the dashboard cannot run.
What does structured content actually do?
Structure earned its keep, just not in the way most guides claim. Pages with headings, lists, and tables generated 0.50 more citation markers per answer than plain-text pages (95% CI: 0.20 to 0.84), a real and meaningful lift in how much a cited page gets quoted.
But the binary question, whether a page gets cited at all, moved less. Structured formatting produced only a 4.5 percentage point increase in citation incidence (95% CI: -1.4 to +10.4, p=.168), not a statistically reliable result. Structure appears to help a page that is already in the running get quoted more thoroughly. It does not reliably decide whether that page makes the cut in the first place.
Why do identical prompts produce different citations?
The study's most uncomfortable finding has nothing to do with content or rank. The researchers reran 30 frozen query families through fresh decoding with identical inputs, and the binary citation decision flipped in 15% of cases, roughly one in seven. About 45% of the variance in which source got credited traced to randomness in the model's own decoding, not to any difference between the pages.
That is a meaningful share of noise sitting inside a metric marketers treat as a clean signal. A single citation check, run once, is measuring a coin with a documented wobble.
Who does this affect?
Teams making decisions off thin evidence carry the most exposure here.
- SEO and GEO teams chasing position 1 as the primary lever for AI citations
- Content teams treating formatting as a guaranteed citation trigger rather than a quoting multiplier
- Publishers and agencies benchmarking citation share off a single test run
- In-house marketing teams without the infrastructure to test at scale or rerun checks over time
How should you respond?
Stop treating a single citation check as a verdict. With a 15% flip rate on identical inputs, one-off audits are noise dressed as signal. Measurement needs repetition: track a page's citation presence across multiple runs and multiple engines, ChatGPT, Perplexity, Gemini, AI Overviews, before drawing a conclusion about whether it is working.
Stop optimizing for rank as if it were the cause. Spend the effort instead on the things the raw correlation was actually picking up: topical authority, tight relevance to the question being asked, and content depth an agent can lift from cleanly. Formatting still matters, use headings, lists, and tables, but treat it as a way to earn a fuller quote once you're in the mix, not a guarantee of getting in.
This is the same case for scale and freshness over one-off tactics that shows up everywhere else in AI search visibility. LeadHaste went from zero to 1 million impressions and more than 200 AI citations in four months by publishing at volume and keeping content current, not by chasing individual ranking positions. A citation system that quotes probabilistically, with real decoding noise baked in, rewards consistent coverage over a category more than it rewards any single optimized page.
Source: Selvam & Ghosh, CiteChoice preprint, arXiv, 2026-09-14
Key takeaways
- Raw data showed an 85.1% to 42.8% citation gap between rank 1 and rank 5 pages, a 42.3 point spread.
- A controlled test that actually swapped page positions found the real causal effect was only 7.9 points.
- A replication test found no measurable position effect at all, statistically indistinguishable from zero.
- Structured content (headings, lists, tables) added 0.50 more citations per answer but did not reliably change whether a page got cited at all.
- Rerunning identical queries flipped the citation decision 15% of the time, with 45% of outcome variance traced to model randomness.
- Single-run citation checks are unreliable; citation share needs to be measured across repeated tests and multiple engines.
Omnicite Editorial. "AI Citation Behavior: How Search Agents Pick Sources" The Citation Report, Omnicite. https://omnicite.co/blog/how-do-ai-search-agents-decide-which-pages-to-ci/
Sources
Source: arXiv (Selvam & Ghosh, CiteChoice preprint)
Raw citation rate gap between rank 1 and rank 5 pages was 42.3 points, and the controlled causal test found only a 7.9 point effect, with a replication test finding no measurable effect; structured formatting added 0.50 citations per answer; identical reruns flipped 15% of citation decisions. arXiv (Selvam & Ghosh, CiteChoice preprint), 2026-09-14
Source: atozseo.in
News summary and interpretation of the CiteChoice study's findings on AI search citation distribution atozseo.in, 2026-09-17
Frequently asked questions
What is the CiteChoice study?
CiteChoice is a preprint by Sriram Selvam and Anneswa Ghosh, published on arXiv on September 14, 2026, that causally tests how a GPT-5.4 search agent distributes citations across pages that support the same fact.
Does ranking position still affect AI citations?
Some, but far less than raw data suggests. The controlled test found a 7.9 percentage point effect from position, and a replication test found no measurable effect, well below the 42.3 point gap the uncontrolled data implied.
Does content formatting help you get cited by AI search agents?
Structured content (headings, lists, tables) increased how many citation markers a page received once it was already part of the answer, by 0.50 per answer on average. It did not reliably increase whether the page got cited at all.
Why did the same query produce different citations on separate runs?
The study found that rerunning identical queries with fresh decoding flipped the citation decision in 15% of cases. Roughly 45% of that variance came from randomness in how the model generates its answer, not from any difference in the source pages.
Who does this study affect most?
SEO and GEO teams optimizing primarily for rank position, publishers benchmarking citation share off a single test, and anyone treating one AI citation check as a reliable measurement.
What should teams do differently based on this study?
Track citation presence across repeated runs and multiple engines instead of a single check, and prioritize topical authority and content depth over chasing rank position as the primary lever.