[The Engines]

How Can Brands Navigate the Complexities of AI Citation Behavior?

A new controlled audit of AI search agents found that raw position-to-citation correlations mostly evaporate under testing, while structured content still earns a real, measurable citation gain.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

A September 2026 controlled study of AI search agents found that the popular claim 'top position gets cited twice as often' collapses once researchers isolate position from other variables, while rendering a fact as a stat, list, or table produces a real citation gain of about half a citation per answer. The fix is not to chase rank. It is to make your page the one carrying the structured, sourced version of the claim.

What did the new study change about AI citation behavior?

It changed the confidence with which anyone can point to search-result position and call it the reason a page gets cited. A preprint titled CITECHOICE, posted to arXiv on September 14, 2026 by researchers Sriram Selvam and Anneswa Ghosh, ran a causal audit of how an AI search agent (GPT-5.4 paired with the Exa search provider) allocates citations when two sources support the same fact.

The team pulled 129 everyday-query transcripts and isolated 113 document pairs where both sources independently verified the same pre-specified claim. They then replayed each transcript twice, swapping the order of the pair and rewriting one target source as either structured (stats, bullets, a short table) or plain prose, while holding everything else fixed. That design lets the study separate correlation from cause, which is the part most GEO commentary skips.

The raw numbers looked dramatic before controls: sources in the top position were cited in 85.1% of transcripts, versus 42.8% for sources ranked fifth, a 42.3-point gap, according to AtoZ SEO's summary of the working paper. Once the researchers controlled for other variables, that gap fell to 7.9 percentage points in one test and to a statistically flat 0.0 points (95% confidence interval, -5.4 to +5.4) in a second. Position, it turns out, was mostly standing in for formatting.

Rewriting a source in structured form (the CITECHOICE authors call it 'structured rendering') raised the target's citation count by +0.50 citations per answer, with a 95% confidence interval of +0.20 to +0.84 and a Holm-adjusted p-value of .033. Crucially, that gain did not come by stealing credit from a competing source. It concentrated citation credit onto the structured page without reducing total citations across the answer. A third finding is a caution for anyone running single-prompt spot checks: citation decisions flipped in 15% of identical reruns of the same query, and the researchers estimate roughly 45% of the variation they saw came from model randomness rather than anything about the page itself.

Who does this change affect?

Two groups feel this most directly. B2B SaaS and tech growth teams that have been told to 'win position one' on comparison and 'best tool' prompts now have evidence that position is a weaker lever than they assumed, and that the structured asset on the page (a comparison table, a named stat, a dated data point) is doing more of the real work.

Local, multi-location, and service businesses face a related problem from the other direction. A business that ranks well for '[service] in [city]' but presents that information as a wall of prose is leaving citation credit on the table that a structured competitor could pick up, even from a lower-ranked page.

It also affects anyone selling GEO tactics as position-gaming. If a vendor's pitch rests on 'we get you to position one and the citations follow,' this study is the first controlled evidence that the correlation behind that pitch is weaker than the industry has assumed. That does not mean position is worthless. It means position without a structured, sourced claim on the page is doing less than most GEO decks imply.

Before and after: what the CITECHOICE audit changed about AI citation behavior (posted September 14, 2026)
MetricBefore: raw assumptionAfter: controlled findingWhat to do now
Position effectTop-ranked sources cited in 85.1% of transcripts vs 42.8% for fifth place, a 42.3-point gapGap shrank to 7.9 points in one controlled test and to 0.0 points (95% CI -5.4 to +5.4) in a secondStop treating rank alone as a citation lever. Pair it with a structured, sourced claim on the page.
Formatting effectAssumed neutral, rarely isolated from other variables in prior testingStructured rendering added +0.50 citations per answer (95% CI +0.20 to +0.84, p=.033) without reducing competitors' creditAdd a dated stat, table, or data point to every page competing for a shared claim.
Test reliabilityA single prompt-and-response check was treated as representativeCitation decisions flipped in 15% of identical reruns; about 45% of variance traced to model randomnessRerun the same prompt multiple times, across more than one AI engine, before concluding a page is or is not cited.

How should brands respond to the new citation research?

Respond by moving budget and attention from rank-chasing to asset-building, and by testing claims more than once before acting on them. Three concrete moves follow directly from the study's findings.

First, put a real citable asset on every page competing for a claim: a dated statistic, an original data point, or a comparison table, each with a named, linked source. The +0.50 citation gain in CITECHOICE came specifically from structured rendering of a verifiable fact, not from prose that merely mentions the fact.

Second, stop treating a single query and a single answer as proof a page is or is not being cited. With a 15% flip rate on identical reruns, one test tells you little. Run the same prompt several times, across more than one AI engine (ChatGPT, Perplexity, Gemini, and AI Overviews all allocate citations differently), before drawing a conclusion about a page's citation performance.

Third, keep rank as one input, not the strategy. A page still has to be found and considered before it can be cited, so ranking and indexing work still matters. What changes is the story you tell a client or a stakeholder about why a page got cited: the honest answer, per this study, leans toward the structure and sourcing of the content on the page rather than its position in a results list.

What should brands watch for before overreacting?

CITECHOICE is a preprint that has not yet been peer reviewed, and the authors tested one model and one search provider, GPT-5.4 with Exa, replayed offline instead of against live, editable webpages. The findings are a strong, controlled signal, not a settled law for every AI engine.

The caution cuts both ways. It is a mistake to conclude position never matters, and it is equally a mistake to conclude formatting is a trick that reliably wins a citation. The honest framing, consistent with the study's own recommendation, is to test in aggregate, across repeated runs and more than one provider, instead of acting on a single query result in either direction.

Key takeaways

  • A September 2026 controlled study found that the raw gap between top-position and fifth-position citation rates (85.1% vs 42.8%) mostly disappears once other variables are isolated.
  • Rewriting a source's content as structured stats, lists, or tables raised its citation count by +0.50 per answer without pulling credit away from competing sources.
  • Citation decisions flipped in 15% of identical reruns of the same query, so a single spot-check is not reliable evidence either way.
  • B2B SaaS teams chasing 'position one' on comparison prompts and local or service businesses presenting facts as prose both have reason to rebuild pages around a structured, sourced claim.
  • The study is a preprint, tested on one model and one search provider, so treat it as a strong signal to test against, not a settled rule for every AI engine.
  • The practical response is to add a citable asset to every competing page and to test citation outcomes across repeated runs and multiple engines before acting.

Omnicite Editorial. "AI Citation Behavior: What Really Changed" The Citation Report, Omnicite. https://omnicite.co/blog/how-can-brands-navigate-the-complexities-of-ai-c/

Sources

Source: arXiv (Selvam, S. and Ghosh, A.)

CITECHOICE study abstract, methodology, and the +0.50 citations per answer structured-rendering finding with confidence intervals arXiv (Selvam, S. and Ghosh, A.), 2026-09-14

Source: AtoZ SEO

Raw, uncontrolled citation rates by search-result position (85.1% vs 42.8%) and the 15% rerun flip rate reported from the working paper AtoZ SEO, 2026-09-17

Source: arXiv (Aggarwal, P. et al., GEO: Generative Engine Optimization, KDD 2024)

Foundational finding that optimizing content for generative engines, including adding statistics and citations, can boost visibility in AI answers by up to 40% arXiv (Aggarwal, P. et al., GEO: Generative Engine Optimization, KDD 2024), 2024-06-28

Frequently asked questions

What is the CITECHOICE study and who wrote it?

CITECHOICE is a preprint audit of how AI search agents allocate citations among sources that support the same fact. It was posted to arXiv on September 14, 2026 by researchers Sriram Selvam and Anneswa Ghosh, and tested GPT-5.4 paired with the Exa search provider.

Does search-result position still matter for AI citations?

Some, but far less than raw numbers suggest. The uncontrolled gap between top and fifth position (85.1% vs 42.8% citation rate) shrank to 7.9 percentage points in one controlled test and to a statistically flat 0.0 points in a second, meaning position alone is a weak, inconsistent lever.

Does formatting content as a stat or table actually increase citations?

Yes, in this study. Structured rendering of a fact (as stats, bullets, or a table) raised the source's citation count by +0.50 citations per answer (95% CI +0.20 to +0.84) without reducing the citations any competing source received.

Why did citation decisions change across identical test runs?

The study reran identical queries and found citation decisions flipped 15% of the time, with the researchers estimating about 45% of that variation came from model randomness rather than anything specific to the page being cited.

Is this study conclusive enough to change a content strategy?

It is a controlled signal worth acting on, not a settled law. It is a preprint tested on one model and one search provider, replayed offline rather than against live pages, so pair its direction (build structured, sourced assets) with your own repeated testing across engines.

What should a brand do differently starting now?

Put a dated stat, data point, or comparison table on every page competing for a shared claim, and stop drawing conclusions from a single prompt-and-response check. Test the same prompt multiple times and across more than one AI engine before deciding a page is or is not being cited.