[The Engines]

How Can Brands Navigate the Randomness in AI Citation?

A September 2026 study found AI answer engines flip citation choices 15% of the time on identical reruns, and raw position data oversold its own effect by more than 5x. Here is what changed and how to adjust your tracking.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

A September 14, 2026 study called CITECHOICE found that AI answer engines flip their citation choice in 15% of identical reruns, and that the raw correlation between position and citation (42.3 points) overstated the true, isolated effect (7.9 points) by more than five times. Brands should stop judging citation from one-off checks and start tracking citation rate across repeated prompts and dates.

What did the new AI citation study actually find?

A paper published September 14, 2026 on arXiv, called CITECHOICE, ran a controlled test of how AI answer engines choose which source to cite when several documents say the same thing. Researchers Sriram Selvam and Anneswa Ghosh built the study around 130 questions, then isolated 113 verified pairs of competing sources, with 103 of those pairs confirmed by blinded human review. The change is not that AI systems started citing differently overnight. The change is that someone finally separated correlation from cause in how those citations get handed out.

Earlier tracking tools mostly reported what lines up with citation: pages ranked first got cited far more often than pages ranked fifth, so vendors told brands to chase position one. CITECHOICE tested that claim directly. Researchers swapped document position while holding the content fixed, then measured what actually moved. The isolated effect turned out to be a fraction of the raw correlation, and that gap is the real story here.

The dataset itself is worth sitting with for a moment. A hundred or so verified pairs is not a huge sample by web-scale standards, but the study's design, controlled swaps instead of passive observation, is what makes the numbers trustworthy. Most citation research to date has been observational: scrape a batch of AI answers, count what got cited, publish a correlation. CITECHOICE instead manipulated one variable at a time and measured what changed, which is why its conclusions carry more weight than another leaderboard of 'top cited domains.'

Reaction pieces have already started circulating. A write-up on AtoZ SEO published September 17, 2026 called the paper a look inside the black box of how citation gets distributed, and warned that raw statistical correlations between a page's placement and its citation rate can be deeply misleading. That warning is worth taking seriously if your reporting stack treats a citation count as a fixed score instead of a sampled estimate.

This matters now because the study lands right as more brands build whole reporting programs around raw citation counts. A methodology paper that quantifies the noise in those counts is not academic housekeeping. It changes what a reasonable citation report is allowed to claim.

Why do raw citation metrics lie?

Raw metrics lie because they mix two different things: what tends to sit next to a citation, and what actually causes one. In the CITECHOICE data, top-ranked sources were cited 85.1% of the time, against 42.8% for fifth-ranked sources, a 42.3 percentage point gap. That number, on its own, looks like proof that ranking first is most of the game.

It is not. When researchers experimentally swapped a document's position and held everything else constant, the isolated position effect dropped to 7.9 percentage points. A secondary test found an effect of 0.0 points, with a confidence interval spanning -5.4 to +5.4. Position correlates with citation mostly because stronger content tends to rank higher for reasons that have nothing to do with rank itself. Strip that confound out and the position lever shrinks by more than five times.

Formatting shows a similar, smaller pattern. Rewriting a page into a more structured layout, with clear headers and direct answers, added roughly 0.50 citation markers per answer (95% confidence interval: 0.20 to 0.84) and raised citation probability by about 4.5 percentage points. That lift is real and usable. It is not the guaranteed citation trigger that some content tools imply when they sell 'AI-optimized formatting' as a fix for everything.

None of this means quality stops mattering. It means quality earns citation through a noisier channel than a simple leaderboard suggests. Coverage, freshness, and clarity still separate a page that gets cited from one that doesn't; the study just proves that position and formatting alone explain less of the outcome than raw dashboards imply, and that no amount of gaming a single lever produces a guaranteed result.

CITECHOICE (September 14, 2026): raw correlation vs. isolated causal effect
SignalRaw / observed readingControlled / causal reading
Position (rank 1 vs. rank 5)42.3 point gap in citation rate (85.1% vs. 42.8%)7.9 point effect when position alone was swapped; 0.0 points in a secondary test
Structured formattingOften marketed as a near-guaranteed citation trigger+0.50 citation markers per answer; +4.5 points of citation probability
Citation decision stabilityTreated as a fixed, repeatable outcome15% of citation choices flipped on identical reruns

Who does this affect?

The finding lands hardest on two groups already tracked in most citation dashboards: B2B SaaS and growth teams watching whether ChatGPT recommends them instead of a competitor for a category prompt, and local or service businesses checking whether they show up when someone asks an AI assistant for the best option in their city. Both groups tend to check citation the same flawed way, running a prompt once and treating the screenshot as a verdict.

The finding also affects anyone building or buying a citation-tracking tool. If a dashboard reports that a brand ranks third and gets cited 30% of the time, without accounting for the noise CITECHOICE measured, that dashboard is presenting a snapshot as if it were a stable trend.

There is a second layer worth flagging. A July 2026 Search Engine Journal analysis of more than 150,000 citations found 96% pointed to third-party sources, Reddit, YouTube, forums, and industry publications, rather than to the brand's own site. Combine that with the noise CITECHOICE measured and the picture gets harder: brands are being judged by citation patterns on pages they mostly don't control, and those patterns shift even when nothing on the page changes.

For local and service businesses, the stakes are practical rather than academic. A single prompt like 'best plumber in Austin' can return a different cited business on back-to-back runs with nothing about either business changing in between. Judging a local campaign off one afternoon of manual prompt checks produces a misleading read almost every time, in either direction.

How random is AI citation, really?

Random enough to break a single-check methodology. CITECHOICE reran identical prompts against identical documents and found the citation decision flipped in 15% of cases. Roughly 45% of the variation the researchers observed across outcomes traced back to model randomness rather than any measurable difference in content, position, or formatting.

That does not make citation meaningless. It makes a one-time check meaningless. Checking once, for a single prompt, on a single day, is closer to a coin flip with weighted odds than to a measurement. The weighting is real and worth optimizing. The single flip is not a verdict.

This is exactly why the industry has been moving toward Citation Share, the percentage of relevant AI answers in a category that cite a given brand, measured across many prompts and repeated over time, instead of a single-answer check. Averaging across enough prompts and enough reruns is the only way to see past a 15% flip rate and find the real trend underneath it.

How should brands respond to citation randomness?

Respond by changing what counts as a measurement, not by abandoning the metric. Five adjustments follow directly from the CITECHOICE numbers.

  1. Stop drawing conclusions from a single prompt check. Rerun the same prompt across multiple sessions and dates before deciding whether you were or were not cited.
  2. Track citation rate over a rolling window instead of a single day's citation count, since the same content can flip outcomes 15% of the time with nothing changed.
  3. Keep investing in structured, clearly organized pages. The lift is real, about 4.5 percentage points of citation probability, even though it is modest rather than decisive.
  4. Deprioritize chasing rank one as a whole strategy. The isolated position effect (7.9 points) is a fraction of the raw correlation (42.3 points) that made position look dominant.
  5. Budget for volatility in citation reporting the way you would for any noisy marketing channel, and treat week-over-week swings as expected rather than alarming.

What should you stop doing?

Stop treating a missed citation on one prompt as proof that a page failed. Don't buy into 'guaranteed AI citation' claims tied to a single formatting change or a single ranking position, since the paper's own confidence intervals include zero effect in more than one test. And stop reading raw backlink-style or rank-style correlations as if they were causal drivers of citation. CITECHOICE is direct evidence that they usually are not.

The honest framing is closer to early web analytics before sampling and attribution matured: the signal is there, it is just noisier than the first-generation dashboards admitted. Brands that build repeated measurement into their process, checking citation across many prompts and many reruns, will read that signal correctly. The ones still reacting to single-day snapshots will keep mistaking noise for a trend, in both directions.

Consider the underlying incentive shift, too. If pure ranking position mattered as much as raw correlation implied, the winning strategy would be gaming a single leaderboard. CITECHOICE suggests otherwise: the bigger, more reliable lever is coverage across many relevant prompts, because averaging out a 15% noise floor takes volume and consistency, not a single trick applied to a single page. That is closer to how Citation Engineering approaches like Omnicite's treat the problem: publish enough well-sourced, well-structured coverage across a category that the noise washes out and the trend line still points up. There is no page two in an AI answer, and there is no reliable single measurement either. Both are reasons to build for the average outcome across a lot of prompts, not the result of the one you happened to check this morning.

Key takeaways

  • A September 2026 study (CITECHOICE) found AI citation choices flip in 15% of identical reruns, so a single citation check is not a reliable measurement.
  • Raw position-to-citation correlation (42.3 percentage points) overstated the true, isolated effect of position (7.9 points) by more than five times.
  • Structured formatting still helps, adding roughly 4.5 percentage points of citation probability, but it is not the guaranteed trigger some vendors claim.
  • Most AI citations still land on third-party pages such as Reddit, YouTube, and forums instead of brand-owned domains, per a July 2026 Search Engine Journal analysis.
  • Brands should track citation rate across repeated prompts and dates, not one-off snapshots, to separate real signal from model noise.
  • Treat AI citation like any noisy channel: budget for swings, sample repeatedly, and read the trend line instead of a single data point.

Omnicite Editorial. "AI Citation Randomness: What Brands Should Know" The Citation Report, Omnicite. https://omnicite.co/blog/how-can-brands-navigate-the-randomness-in-ai-cit/

Sources

Source: arXiv (Sriram Selvam and Anneswa Ghosh, CITECHOICE)

Citation decisions flipped in 15% of identical reruns; raw position correlation (42.3 points) exceeded the isolated causal effect (7.9 points); structured formatting added about 4.5 percentage points of citation probability. arXiv (Sriram Selvam and Anneswa Ghosh, CITECHOICE), 2026-09-14

Source: AtoZ SEO

News analysis summarizing the CITECHOICE findings on citation distribution and why raw metrics can be misleading. AtoZ SEO, 2026-09-17

Source: Search Engine Journal

96% of AI search citations point to third-party sources such as Reddit, YouTube, and forums rather than brand-owned domains. Search Engine Journal, 2026-07-07

Frequently asked questions

What is the CITECHOICE study?

CITECHOICE is a research paper by Sriram Selvam and Anneswa Ghosh, published September 14, 2026 on arXiv, that ran 130 queries through an agentic search pipeline and used controlled experiments, swapping document position and formatting, to isolate what actually causes an AI answer engine to cite one source over another.

Does this mean AI citation is completely random?

No. The study found that about 45% of the variation in outcomes traced to model randomness, which leaves roughly 55% explained by real differences in content, position, and formatting. Citation is noisy, not meaningless.

Should brands stop caring about ranking position for AI visibility?

No, but the position lever is smaller than raw data suggests. The isolated causal effect of moving from rank five to rank one was 7.9 percentage points, well below the 42.3 point gap that raw correlation implied.

How much does structuring content actually help with AI citation?

CITECHOICE measured a real but modest lift: roughly 0.50 more citation markers per answer and about 4.5 percentage points higher citation probability for structured, clearly organized pages. It is a genuine lever, not a guarantee.

How often should a brand check whether AI answer engines cite it?

More than once. Because citation decisions flipped in 15% of identical reruns in the CITECHOICE study, a single prompt check on a single day is not a reliable read. Track citation rate across repeated prompts and dates instead.

What is Citation Share and how does it relate to this study?

Citation Share is the percentage of relevant AI answers in a category that cite a given brand, tracked across many prompts and repeated over time. Measuring it that way is the practical fix for the 15% single-query flip rate CITECHOICE documented, since averaging across volume turns a noisy signal into a usable trend.