[The Engines]

How ChatGPT's Browsing Tool Decides Which Pages to Open and Summarize

ChatGPT does not fetch pages at random. It runs every candidate URL through a set of legibility, trust, and freshness checks before it ever gets quoted in an answer.

The short answer

ChatGPT's browsing tool favors pages whose titles and URLs closely match the exact sub-query it generated, pulls most of its cited pages from its own core search index rather than forums or video, and skews toward content that is roughly a year old or newer. You can shift the odds in your favor with legible URLs, sub-query-matching titles, and by keeping OAI-SearchBot unblocked, but none of it guarantees a citation. See geo for the broader discipline this sits inside.

How does ChatGPT decide which URLs to visit when browsing?

ChatGPT does not browse the open web on a whim. When a prompt calls for current or specific information, ChatGPT rewrites it into one or more targeted search queries, sends those to its search stack, retrieves a set of candidate pages, and only then decides which ones are worth opening and quoting. OpenAI described this pipeline when it launched the feature: ChatGPT 'can now search the web in a much better way than before' and returns 'fast, timely answers with links to relevant web sources' (OpenAI, 2024-10-31).

Three separate OpenAI agents sit behind this process, and they do different jobs. GPTBot crawls the web to help train future models. OAI-SearchBot crawls and indexes pages specifically to power ChatGPT's search and citation features. ChatGPT-User fires in response to a live user request, fetching a specific page in real time to answer that one question. OpenAI's own crawler documentation is explicit that these three are controlled independently in robots.txt, so a site can allow OAI-SearchBot for search visibility while disallowing GPTBot for training, as reported after OpenAI's late-2025 documentation revision (PPC Land, 2025-12-09).

Retrieval is only the first cut. Ahrefs' analysis of 1.4 million ChatGPT prompts found that roughly half of all retrieved pages ever make it into a visible citation. That gap, between what ChatGPT looks at and what it credits, is where the real selection happens, and it is the part worth reverse engineering (Ahrefs, 2026-04-15).

What makes a page get opened by ChatGPT's browsing tool?

The single strongest signal is whether a page's title and URL match the specific sub-query ChatGPT generated, not just the broad topic of the original prompt. Ahrefs measured this directly: cited pages averaged 0.602 semantic similarity to the prompt, versus 0.484 for pages that were retrieved but never cited, and the gap widened further against ChatGPT's internal sub-queries (Ahrefs, 2026-04-15).

Where a page comes from matters almost as much as what it says. In the same study, pages surfaced through ChatGPT's core search index were cited 88.46% of the time. Pages surfaced through news sources were cited 12.01% of the time. Reddit results were cited just 1.93% of the time despite being retrieved constantly, and YouTube results 0.51%. ChatGPT appears to lean on forum and video content to build context and gauge consensus, then routes the actual citation to a page it treats as a more authoritative primary source.

Recency plays a smaller but measurable role. The average cited page in the dataset was about 500 days old, and within news results specifically, cited articles skewed noticeably younger than the ones retrieved but passed over. Freshness is not a hard cutoff, it is a tiebreaker.

URL legibility rounds out the picture. Pages with natural-language URL slugs saw an 89.78% citation rate compared with 81.11% for pages with opaque, parameterized, or ID-based URLs, a gap Ahrefs attributes to how much easier a plain-language slug is to match against a sub-query (Ahrefs, 2026-04-15).

  1. Query-to-title/URL match: does the page's title or URL echo the exact sub-query, not just the topic
  2. Retrieval channel: search index (88.46% cited) versus news (12.01%), Reddit (1.93%), YouTube (0.51%)
  3. Recency: cited pages average roughly 500 days old, news skews younger
  4. URL legibility: natural-language slugs cited more often than opaque URLs
  5. Crawler permission: OAI-SearchBot must be allowed in robots.txt or the page cannot be cited at all
Signals ChatGPT's browsing tool checks before opening and citing a page
SignalWhat ChatGPT is checkingWhy it mattersEvidence
Query-to-title/URL matchDoes the title or URL align with the specific sub-query, not just the broad topicCited pages average 0.602 similarity to the sub-query versus 0.484 for retrieved-but-ignored pagesAhrefs, 1.4M-prompt study, 2026-04-15
Retrieval channelDid the page surface via the core search index, or via news, Reddit, YouTube, or academic sourcesSearch-index results are cited 88.46% of the time; Reddit 1.93%, YouTube 0.51%Ahrefs, 1.4M-prompt study, 2026-04-15
RecencyHow old the page is relative to competing results for the same sub-queryAverage cited page is about 500 days old; cited news skews younger than uncited newsAhrefs, 1.4M-prompt study, 2026-04-15
URL legibilityIs the URL a natural-language slug or an opaque, parameterized pathNatural-language slugs see an 89.78% citation rate versus 81.11% for opaque URLsAhrefs, 1.4M-prompt study, 2026-04-15
Crawler permissionDoes robots.txt allow OAI-SearchBot, independent of whether GPTBot is allowedBlocking OAI-SearchBot removes a page from ChatGPT Search citations entirely, regardless of content qualityOpenAI crawler documentation, reported by PPC Land, 2025-12-09

Does ChatGPT's browsing tool still respect robots.txt?

Partly, and the answer changed recently. GPTBot and OAI-SearchBot still honor robots.txt disallow rules, and OpenAI's documentation confirms the two can be blocked or allowed independently of each other. Blocking GPTBot keeps a site out of model training data without removing it from ChatGPT Search citations, and vice versa (PPC Land, 2025-12-09, reporting on OpenAI's crawler documentation).

ChatGPT-User is the exception. OpenAI's revised documentation frames ChatGPT-User's fetches as user-initiated rather than autonomous crawling, and states plainly that because these actions are initiated by a user, robots.txt rules may not apply to them. In practice, a page can be blocked from OAI-SearchBot's indexing crawl and still get opened and read in response to an individual user's live question.

This is where the mechanics connect to the wider answer engine landscape: the crawler that decides whether you show up in search results is not the same crawler that decides whether a single user's question triggers a live fetch of your page. Treating 'block the bots' as one setting misses that split.

Can I influence which pages ChatGPT fetches?

Yes, on the margins that are actually yours to control, not by manipulating the model. The honest framing is that you can raise the probability a page gets opened and cited; nothing in OpenAI's documentation or Ahrefs' data supports the idea of a guaranteed citation.

The concrete levers line up with the signals above: keep OAI-SearchBot allowed in robots.txt so the page is eligible for citation at all, write titles and URL slugs in the plain language a person would actually ask ChatGPT, and keep content current enough that it does not read as stale next to a fresher competing source. None of these are exotic. They are closer to disciplined publishing hygiene than to any kind of trick.

What they cannot do is substitute for coverage and authority. A perfectly worded title on a thin, one-off page will still lose to a page that thoroughly answers the sub-query ChatGPT actually generated. The signals reward specificity and depth, not phrasing alone.

What does this mean for how you publish?

It means the unit of optimization is the sub-query, not the keyword. A page built around one broad topic will get retrieved for many related prompts and cited for almost none of them, because it rarely matches the narrow, rewritten sub-query ChatGPT is actually scoring against.

It also means volume and freshness compound. Publishing thoroughly, at a pace that keeps content inside the recency window Ahrefs measured, across the specific sub-questions a category generates, is what moves a domain from occasionally retrieved to consistently cited. That is the operating premise behind Citation Share as a metric in the first place: it tracks how often you are the page that gets opened and quoted, not just the page that shows up somewhere in the retrieval set.

None of this requires guesswork if you track it. Watching which of your pages actually get cited across ChatGPT, Perplexity, and Gemini, and which ones only get retrieved and ignored, tells you exactly where your titles, structure, or freshness are falling short of the sub-queries people are actually asking.

Key takeaways

  • ChatGPT rewrites prompts into targeted sub-queries first, then decides which retrieved pages to open and cite, and only about half of retrieved pages ever get cited.
  • The strongest citation signal is title and URL alignment to the exact sub-query, not the broad keyword.
  • Where a page comes from matters: core search-index results are cited far more often than Reddit or YouTube results, even when those sources are retrieved constantly for context.
  • GPTBot, OAI-SearchBot, and ChatGPT-User are controlled independently, and OpenAI's own documentation now states ChatGPT-User's live fetches may not follow robots.txt at all.
  • Recency and plain-language URL slugs are measurable, real tiebreakers, not folklore.
  • None of these signals can be gamed into a guaranteed citation; they shift probability, and depth and coverage still decide the outcome.

Omnicite Editorial. "ChatGPT Browsing: How It Picks Pages to Open" The Citation Report, Omnicite. https://omnicite.co/blog/chatgpt-browsing-tool-page-selection/

Sources

ChatGPT rewrites a prompt into one or more targeted search queries and returns links to relevant web sources OpenAI, 2024-10-31

Across 1.4 million ChatGPT prompts, core search-index results were cited 88.46% of the time versus 1.93% for Reddit and 0.51% for YouTube, and cited pages averaged 0.602 title/URL similarity to the sub-query versus 0.484 for uncited pages Ahrefs, 2026-04-15

Natural-language URL slugs were cited 89.78% of the time versus 81.11% for less descriptive URLs, and cited pages averaged roughly 500 days old Search Engine Journal, 2026-04-16

GPTBot, OAI-SearchBot, and ChatGPT-User are controlled independently in robots.txt, and ChatGPT-User's user-initiated fetches may not follow robots.txt rules at all PPC Land, 2025-12-09

Frequently asked questions

Does ChatGPT crawl the entire internet in real time before answering?

No. ChatGPT rewrites a prompt into one or more targeted search queries, retrieves a limited set of candidate pages for those queries, and only opens or cites a subset of what it retrieves.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot crawls content for model training. OAI-SearchBot crawls and indexes content specifically for ChatGPT's search and citation features. OpenAI's documentation confirms the two are controlled independently in robots.txt.

If I block GPTBot, does that stop ChatGPT from citing my site?

No. Blocking GPTBot only opts a site out of training data collection. OAI-SearchBot, the crawler tied to search citations, has to be blocked separately to remove a site from ChatGPT Search results.

Does robots.txt still control everything ChatGPT fetches?

Not entirely. OpenAI's revised crawler documentation states that ChatGPT-User, which fires for live, user-triggered fetches rather than automated crawling, may not follow robots.txt rules the way GPTBot and OAI-SearchBot do.

Does ChatGPT prefer citing Reddit and forum discussions?

It retrieves them often but rarely cites them. In Ahrefs' study, Reddit pages were cited only 1.93% of the time despite heavy retrieval, suggesting ChatGPT uses forum content for context rather than as the credited source.

How old can a page be and still get cited?

There is no hard cutoff, but cited pages averaged about 500 days old in Ahrefs' dataset, and cited news articles specifically skewed younger than news articles that were retrieved but not cited.