[The Engines]

How Can You Optimize Content to Increase AI Citation Rates?

A new causal study just showed that the 'rank first, get cited more' theory behind most AI citation advice is mostly a correlation, not a cause. Here is what the controlled test actually found and what to change because of it.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

A controlled study published September 14, 2026 found that source position barely affects AI citation once you control for confounding factors: a raw 42.3 point gap between first and fifth position shrank to 0 to 7.9 points under a controlled swap test. What did move citations, modestly, was structuring pages with headings and lists. Stop chasing position in a single test and start building for extraction and multi-run consistency.

What did the new study actually find?

A controlled study of AI citation behavior, published as a preprint on arXiv on September 14, 2026 by researchers Sriram Selvam and Anneswa Ghosh, tested a claim that has driven two years of GEO advice: that ranking higher in a source list makes an AI answer more likely to cite you. The paper, called CiteChoice, audited roughly 129 real multi-turn search transcripts and isolated 113 matched source pairs, cases where two documents supported the identical fact and only their order or formatting changed between runs.

The raw numbers told a clean story on their own. Pages placed first in the source list got cited 85.1 percent of the time. Pages placed fifth got cited 42.8 percent of the time. That is a 42.3 point gap, exactly the kind of correlation a vendor dashboard would present as proof that position drives citation, and exactly the kind of number that has justified a year of 'get ranked first' advice in the GEO space.

Then the researchers did the part most vendor reports skip: they swapped the positions of matched pairs and reran the test under controlled conditions, holding the rest of the transcript fixed. The isolated position effect collapsed to somewhere between 0.0 and 7.9 percentage points depending on the test set, and several runs lost statistical significance entirely. Most of that 42-point gap was never really about slot order. It reflected which pages tend to land first for reasons unrelated to position: relevance to the query, specificity of the claim, and how easily a model could extract a fact from the page.

That distinction between correlation and cause is the whole story. A page that already answers a question clearly and specifically tends to both rank first and get cited most. Move it to fifth and it keeps most of its citation rate. Move a weaker page to first and its citation rate barely rises. Position was riding on top of quality the entire time, not driving it.

One variable did produce a real, independent effect: formatting. Rewriting a page with headings and lists produced an average of 0.50 additional citation markers per answer compared with a plain paragraph version. That effect is real, but the researchers are explicit that it redistributes citation credit within an answer instead of proving that formatting alone wins more total citations across the board.

A third finding may matter more than the headline result. When the team reran identical queries against the same model, reported in coverage as GPT-5.4, the decision to cite or not cite the same target page flipped in 15 percent of cases, with roughly 45 percent of that swing attributable to plain model randomness instead of anything on the page. A single test run cannot separate a real effect from noise at that scale, which is exactly why the authors call their own paper 'an attribution-sensitivity warning, not an optimization tactic.'

The timing matters too. GEO has spent roughly two years treating AI citation the way early SEO treated Google's ten blue links: find the ranking factor, push hard on it, repeat. CiteChoice is the first controlled test of that assumption, not another correlational scrape, and it lands at a moment when plenty of budget is already committed to position-chasing tactics that this data says were misdirected.

Who does this change affect?

Anyone running a GEO or AI-visibility program built on single-run rank checks should read this closely. That includes in-house content and SEO teams tracking 'citation rank' the way they once tracked Google rank, and it includes any vendor whose reporting treats raw position-to-citation correlation as proof of a real effect instead of a coincidence worth double-checking.

Comparison content and listicles feel this most directly, since source-order claims get made loudest there. If a content strategy has quietly narrowed to 'get our page listed first in whatever source list the model happens to pull,' this study is the evidence that the strategy rested on a correlation, not a mechanism.

B2B SaaS teams chasing 'best [category] tool' prompts and local and multi-location businesses chasing '[service] in [city]' prompts both sit inside this finding, just from different angles. The SaaS team is more likely to have been optimizing for position in a comparison answer. The local business is more likely to have been optimizing for position in a directory-style list. Neither position bet was as solid as it looked.

It does not touch the logic behind citation-grade content itself. Coverage, freshness, and structural clarity were never about gaming a slot in a list. They were always about giving a model something worth quoting, and that logic survives this study intact, because it was never the thing the study disproved.

Agencies and platforms selling 'citation rank tracking' as a standalone product face the sharpest version of this. A report that shows a client moved from position five to position one, with no controlled test behind it, is now a report that cannot claim credit for a citation increase with any confidence. The honest version of that report ties movement to coverage and structure changes, not slot number.

CiteChoice study: raw signal vs. controlled finding (arXiv, published 2026-09-14)
SignalRaw, uncontrolled readingControlled findingWhat to do about it
Source position (1st vs. 5th)1st cited 85.1% of the time vs. 42.8% for 5th, a 42.3-point gapSwapping matched pairs cut the isolated effect to 0.0 to 7.9 points, often not statistically significantStop optimizing for list position alone; it was never the real driver
Page formatting (headings and lists)Assumed to help only through readabilityStructured rewrites earned 0.50 more citation markers per answer than plain paragraphsStructure every page for extraction: question-shaped headings and scannable lists
Single-run test resultsTreated as a definitive citation verdict15% of citation decisions flipped on an identical rerun; about 45% of the variation was model randomnessNever act on one test; run multi-pass checks before changing strategy

How should you respond to the findings?

Stop treating position in a single test as a verdict. Start treating structure and multi-run consistency as the levers that actually move citation behavior. That is a change in what gets measured and reported, not just in what gets written.

The table below lines up what the raw, uncontrolled signal suggested against what the controlled test found on September 14, 2026, and what each finding means for the way you brief, build, and audit content going forward.

None of this requires new tooling. It requires running the same query more than once before drawing a conclusion, and it requires judging a page on whether it answers a question directly and legibly, not on where it happened to land in one screenshot of one model's output.

Expect the loudest pushback from anyone whose entire pitch was 'rank first, get cited.' That pitch was built on the 42.3 point raw gap, and the raw gap was always the wrong number to build a strategy on.

What should you stop doing?

A few habits built on the raw correlation are worth retiring now that a controlled test exists to check them against.

  1. Stop optimizing content purely to appear first in a source list; the isolated effect of position measured near zero to 7.9 points once confounders were controlled.
  2. Stop reporting a single test run as proof a page 'won' or 'lost' a citation; 15 percent of citation decisions flipped on an identical rerun.
  3. Stop assuming a formatting change boosts total citation volume across the board; the study found it redistributes credit within an answer instead of growing the overall pool.
  4. Stop buying into vendor dashboards that present raw position correlation as a causal driver of citation without a controlled test sitting behind the number.

What should you start doing instead?

The response that actually follows from the data looks less like rank-chasing and more like editorial discipline applied consistently.

  1. Structure every page for extraction: question-shaped headings, scannable lists, and a direct answer near the top, the pattern behind the 0.50-marker gain in the CiteChoice test.
  2. Run multi-pass checks before changing strategy on a page. One test result is noise; a pattern that holds across repeated runs is signal.
  3. Prioritize coverage and freshness over slot position. A page that answers more of the relevant question universe earns citations across more prompts, not just one ranked list.
  4. Keep every statistic and claim dated and sourced. The CiteChoice authors describe their own findings as 'an attribution-sensitivity warning, not an optimization tactic,' and that same caution should apply to any citation claim your team publishes.
  5. Track Answer Presence, how broadly a topic is covered across the relevant question universe, alongside Citation Share, instead of a single position metric for a single prompt.

Does position matter at all now?

A little, in some test sets, but not in the way most GEO advice has assumed. The honest reading is that position is a weak, inconsistent signal riding on top of stronger ones: relevance to the query, structural clarity, and how directly a page answers the question being asked. Chasing the weak signal wastes effort that could go toward the strong ones. That is not a hedge. It is the actual shape of the evidence: weak and inconsistent on its own, strong only when riding on top of relevance and clarity.

That is the same case underneath what this site calls Citation Engineering: building content at the coverage and quality bar an AI system trusts, then tracking Citation Share, the percentage of relevant AI answers in a category that actually cite you, across ChatGPT, Perplexity, Gemini, and Google AI Overviews. LeadHaste's own program moved from zero to 1 million impressions and more than 200 AI citations in four months by publishing at scale on coverage and freshness, not by chasing a slot in a source list. None of that depended on gaming a slot, and none of it needed this study to be true.

What the study does is take away the one shortcut, position gaming, that never worked as advertised in the first place. The teams that were already writing specific, structured, well-sourced answers were never relying on that shortcut, and this is the data that says they were right not to.

Key takeaways

  • The old assumption that ranking first in an AI answer's source list gets you cited more does not hold up under causal testing.
  • A raw 42.3-point citation gap between first and fifth position shrank to 0 to 7.9 points once researchers controlled for confounding factors.
  • Structuring a page with headings and lists produced a real, measurable gain of 0.50 additional citation markers per answer.
  • Citation decisions on an identical query flipped in 15% of reruns, so single-run tests cannot prove a causal effect.
  • GEO teams should shift effort from chasing source-list position to building content that is structurally easy for a model to extract and quote.
  • The findings reinforce the case for citation-grade content built on coverage, freshness, and real sourcing rather than for position-gaming tactics.

Omnicite Editorial. "AI Citation Optimization: What Actually Works" The Citation Report, Omnicite. https://omnicite.co/blog/how-can-you-optimize-content-to-increase-ai-cita/

Sources

Source: arXiv (Selvam and Ghosh, CiteChoice preprint)

CiteChoice causal audit finds the isolated effect of source position on AI citation is near zero to 7.9 points, far below the raw 42.3-point correlation arXiv (Selvam and Ghosh, CiteChoice preprint), 2026-09-14

Source: Search Engine Journal

Analysis of the CiteChoice findings on position, formatting, and citation volatility for search practitioners Search Engine Journal, 2026-09-17

Source: AtoZ SEO

Summary of the CiteChoice study's methodology and why raw citation metrics can be misleading AtoZ SEO, 2026-09-17

Frequently asked questions

What is the CiteChoice study?

CiteChoice is a causal audit of AI citation behavior published as a preprint on arXiv on September 14, 2026 by researchers Sriram Selvam and Anneswa Ghosh. It tested 113 matched source pairs across roughly 129 real multi-turn search transcripts to isolate what actually causes an AI system to cite one source over another.

Does source position still matter for AI citations?

Barely, and inconsistently. The raw data showed a 42.3-point gap between first and fifth position, but once researchers swapped matched pairs and controlled for confounders, the isolated position effect measured between 0.0 and 7.9 percentage points, often without statistical significance.

What formatting actually increases AI citations?

Rewriting a page with headings and lists produced an average of 0.50 additional citation markers per answer compared with plain paragraphs in the CiteChoice test. The study describes this as a redistribution of citation credit rather than a guaranteed increase in total citation volume.

Why do single-run AI citation tests mislead?

Because model outputs are not fully deterministic. When CiteChoice researchers reran identical queries, the decision to cite or not cite the same target page flipped in 15 percent of cases, with about 45 percent of that variation coming from model randomness rather than any real change on the page.

Who published the study and when?

Sriram Selvam and Anneswa Ghosh published CiteChoice as an arXiv preprint on September 14, 2026. Search Engine Journal and other outlets covered the findings starting September 17, 2026.

What should content teams do differently now?

Stop treating source-list position as the primary lever and stop trusting single-run citation checks. Structure content for extraction with clear headings and lists, verify findings across multiple test runs, and keep prioritizing coverage and freshness, the factors most likely to sit behind the position correlation in the first place.