[The Engines]
How Can You Optimize Your Content to Get Cited by AI Search Agents?
A new audit of AI search agents found that most of the citation edge attached to top ranking disappears under controlled testing, and that a single citation test proves almost nothing.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
A new study of AI search agents (arXiv, September 14, 2026) found that the citation gap between a rank 1 and a rank 5 source, 42.3 percentage points in raw data, drops to 7.9 points under controlled reordering and to 0.0 on a held-out repeat. Structured formatting still moves the needle, adding about 0.50 citation markers per answer, but it mostly redistributes credit among sources an AI has already trusted instead of winning a new source admission. Treat any single citation test as noise: 15 percent of citation decisions flipped when researchers reran the exact same prompts.
What did the new study on AI citation behavior actually find?
Ranking position is not the lever GEO and SEO teams have assumed it is. A controlled audit called CITECHOICE, posted to arXiv on September 14, 2026 by researchers Sriram Selvam and Anneswa Ghosh, tested what actually drives citation allocation when an AI search agent has to choose which of several equally supportive sources to quote. The result: the large gap between top-ranked and lower-ranked citation rates comes mostly from correlation, not cause.
The researchers built the test from real search behavior. They ran 129 everyday-query transcripts through a search agent, then isolated 113 document pairs where both sources independently supported the identical fact, confirming 103 of those pairs through blinded human review. From there they ran a hash-verified 2-by-2 replay: swap the order of the paired sources, swap the formatting of one target document between structured and prose, and hold everything else in the transcript fixed. That design separates what position and formatting actually cause from what merely correlates with citation.
Three results survived the controls. Structured rendering concentrates citation credit inside a fixed transcript instead of expanding who gets cited at all. Observational position gaps run far larger than what controlled reordering produces. And citation judgments carry real measurement noise, enough that one test run cannot stand in as evidence of anything.
Why do raw citation position and formatting numbers lie?
Raw numbers overstate what position and formatting can do because they mix cause with everything else correlated with rank one: source authority, topical fit, phrasing, plain luck. In the CITECHOICE data, sources sitting in the top position were cited 85.1 percent of the time against 42.8 percent for fifth position, a 42.3 percentage point gap. That is the number most GEO dashboards would report as proof.
Once the researchers physically swapped which source sat where, holding the rest of the transcript constant, the effect fell to 7.9 percentage points in the main replay and to exactly 0.0 in a held-out repeat, with a 95 percent confidence interval of negative 5.4 to positive 5.4 points. Most of what looked like a position advantage was actually something else about the higher-ranked page, not its rank.
Formatting told a narrower, more honest story. Rewriting a target document into a structured format (headers, lists, direct answers) added about 0.50 citation markers per answer, a real effect (95 percent CI 0.20 to 0.84, Holm-adjusted p equals .033). But the pre-specified test for whether structure got a source cited at all, as opposed to cited more within an answer that already cited it, came back inconclusive at plus 4.5 percentage points (95 percent CI negative 1.4 to positive 10.4). Structure sharpens the credit a trusted source already earns. It does not appear to buy a new source its way into the conversation.
Then there is the noise floor almost nobody accounts for. When the team reran the same frozen transcripts through fresh decoding, 15 percent of binary citation decisions flipped, and roughly 45 percent of the variance in single-run effects came from model randomness rather than anything about the page. A one-off 'I asked ChatGPT and got cited' test is closer to a coin flip than a measurement.
| Signal | Raw / observational reading | Controlled / causal reading | What to do about it |
|---|---|---|---|
| Position (rank 1 vs rank 5) | 42.3 percentage point citation-rate gap (85.1% vs 42.8%) | 7.9 points on the main reorder test, 0.0 points on the held-out repeat | Stop optimizing for rank alone; a lower position is not the handicap raw data suggests |
| Structured formatting | Assumed to broadly increase AI citations | Plus 0.50 citation markers per answer (95% CI 0.20 to 0.84); source admission itself stayed inconclusive at plus 4.5 points | Use structure to sharpen credit for sources already trusted, not as a guaranteed way to get a new page cited |
| Test-to-test consistency | Single-run tests treated as stable results | 15% of citation decisions flipped on identical reruns; about 45% of single-run variance is decoding randomness | Never report or act on a one-off citation test; rerun before concluding anything changed |
Who does this change affect?
Anyone reporting AI citation results from a single test, or selling one as a strategy, is affected.
- SEO and GEO teams who built client reporting around rank-to-citation correlation charts, since the effect behind those charts, once isolated from correlation, is smaller than the raw numbers imply.
- Agencies and vendors pitching structured formatting, schema markup, FAQ blocks, or bulleted answers as a guaranteed way to get cited, when the study shows it redistributes credit rather than reliably winning admission.
- B2B SaaS teams tracking citation share on comparison and best-of prompts, where one favorable screenshot from ChatGPT or Perplexity is not evidence of a trend.
- Local and service businesses evaluating whether a directory listing or landing page earns AI Overviews presence, where one test in one city on one day tells you almost nothing.
How should you respond to the new evidence?
Change what counts as proof before you change what you publish.
- Run every citation test multiple times on the same prompt before drawing a conclusion. A 15 percent flip rate on reruns means one result is an anecdote, not a finding.
- Track both citation count and citation presence, whether the source is cited at all, since the study found these move differently and formatting mainly affects the first.
- Test across more than one model and provider before generalizing a result, because the study's effects were measured on one search agent and one set of questions.
- Weight source admission, being the kind of page an AI agent independently verifies and trusts enough to include, above rank or formatting polish, since admission is the step the study could not show formatting reliably wins.
What should you stop doing right now?
Stop treating rank and formatting as the whole game.
- Stop reporting a single AI citation test as proof a tactic worked. Report a range across repeated runs instead.
- Stop promising clients that reformatting a page into headers and bullet lists will get it cited. It can sharpen credit for a source already in the mix, not guarantee entry.
- Stop building content strategy around beating one competitor's rank position. The position effect the study measured, net of correlation, came in near zero on repeat.
- Stop skipping the question of whether your content gets admitted as a trustworthy source at all. That earlier step, not formatting, is where the real gap between cited and ignored pages likely sits.
Where does this leave AI citation tracking going forward?
It leaves citation tracking where any serious measurement discipline ends up: distrustful of single data points and focused on repeatable signal. The authors themselves describe their formatting result as 'an attribution-sensitivity warning, not an optimization tactic,' a fair summary of the whole paper. Presentation can redistribute credit inside an answer that already trusts your source. It cannot manufacture that trust from nothing.
That distinction is why Omnicite tracks Citation Share, the percentage of relevant AI answers in a category that cite a given brand, across ChatGPT, Perplexity, Gemini and Google AI Overviews, as a trend rather than a single flattering screenshot. The study is a useful check on that discipline: publish content built for coverage and fact accuracy first, use structure to sharpen the credit already earned, and judge the result by what survives a rerun.
Source: arXiv preprint, CITECHOICE (Selvam and Ghosh), 2026-09-14
Key takeaways
- The 42.3 percentage point citation gap between rank 1 and rank 5 sources fell to 7.9 points under controlled reordering, then to 0.0 on a held-out repeat (arXiv, September 14, 2026).
- Structured formatting adds about 0.50 citation markers per answer, but the study could not confirm it reliably wins a new source admission (plus 4.5 points, inconclusive).
- 15 percent of citation decisions flipped when researchers reran identical prompts, and about 45 percent of single-run variance came from model randomness.
- A single AI citation test proves almost nothing; the study's own authors call the formatting effect an attribution-sensitivity warning, not an optimization tactic.
- Source admission, being trusted enough to be cited at all, matters more than rank or formatting once a page is already in the running.
- Citation Share and similar metrics only hold up as a trend measured across repeated prompts and multiple models, not as a one-time screenshot.
Omnicite Editorial. "AI Citation Behavior: How to Respond to the New Study" The Citation Report, Omnicite. https://omnicite.co/blog/how-can-you-optimize-your-content-to-get-cited-b/
Sources
Source: arXiv (preprint)
CITECHOICE causal audit of AI citation position bias, formatting effects, and rerun noise arXiv (preprint), 2026-09-14
Source: Search Engine Journal
Independent reporting and statistical summary of the CITECHOICE findings Search Engine Journal, 2026-09-17
Source: AtoZ SEO
News reaction coverage of the CITECHOICE study that prompted this analysis AtoZ SEO, 2026-09-17
Frequently asked questions
Does search ranking position still affect whether AI search agents cite a page?
Yes, but far less than raw data suggests. The CITECHOICE study found an observational gap of 42.3 percentage points between rank 1 and rank 5 citation rates, but a controlled reordering test cut that to 7.9 points, and a held-out repeat measured 0.0 (arXiv, September 14, 2026).
Does formatting content with headers and lists increase AI citations?
It redistributes citation credit without clearly increasing it. Structured rewrites added about 0.50 citation markers per answer, but the test for whether formatting gets a new source cited at all came back inconclusive.
How reliable is a single AI citation test?
Not reliable. Reruns of the exact same prompts flipped 15 percent of citation decisions, and roughly 45 percent of single-run variance came from model randomness, not the page itself.
Who published the CITECHOICE study and when?
Researchers Sriram Selvam and Anneswa Ghosh posted the preprint to arXiv on September 14, 2026, titled 'CITECHOICE: A Causal Audit of How Document Presentation Redistributes Citation Credit in Agentic Search'.
What should content and SEO teams change based on this study?
Run every citation test multiple times before concluding anything, track both citation count and whether a source is cited at all, and prioritize getting admitted as a trusted source over chasing rank or format alone.
Does this study mean structured content and schema markup are pointless?
No. Structure still sharpens the citation credit a trusted source earns inside an answer. It should not be sold as a guaranteed way to get an untrusted page cited for the first time.