[The Engines]

Can Small Sites Compete in ChatGPT's Search Index?

Small sites are not automatically excluded from ChatGPT Search. The practical challenge is making pages crawlable, easy to summarize and strong enough to earn a citation.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

Small sites can compete in ChatGPT Search. A 2026 reverse-engineering study found no measured serving difference between partner and non-partner pages in an observed in-house retrieval path, but it is not an OpenAI product specification. The response is practical: allow OAI-SearchBot, put the answer near the top of important pages, and measure citations across more than one ChatGPT experience.

What changed in the ChatGPT search index?

The change is that the publisher-deal theory looks less decisive for small-site inclusion than many teams assumed. BotRank reported on August 14, 2026 that Resoneo examined 1,249 ChatGPT answers captured in July and found no measured difference in format, snippet length or freshness between partner and non-partner pages in the observed pipeline it interpreted as OpenAI's in-house index.

That finding is useful because it changes the first diagnostic question. A small publisher, local business or specialist SaaS site should not begin by assuming that a missing commercial distribution agreement is the reason it does not appear in a ChatGPT answer. It should first establish whether OpenAI can crawl the site, whether the page has a direct answer, and whether the page is selected for the prompts that matter.

The evidence has a firm limit. Resoneo's analysis was based on observed network traffic and page-level outputs, not a published OpenAI explanation of ranking or citation selection. BotRank also reported that the relevant pipeline labels stopped appearing around July 21. Treat the result as a dated observation, not a permanent map of how ChatGPT Search works.

OpenAI's own documentation establishes the part that is not speculative: OAI-SearchBot is the crawler used to surface websites in ChatGPT Search. OpenAI says a site that opts out of OAI-SearchBot will not appear in ChatGPT Search answers, although it can still appear as a navigational link. That makes crawl permission a real gate. A publisher deal is not presented in OpenAI's crawler documentation as the universal gate.

The important distinction is between being eligible for retrieval and earning a citation. Eligibility puts a page in contention. Citation depends on the query, the available sources, the answer format and what the system chooses to use. A site can clear the first condition and still be absent from the answer users see.

  1. Before July 21, 2026: Resoneo's observed labels let researchers distinguish an in-house retrieval path from Google scraping, according to BotRank's report.
  2. After July 21, 2026: BotRank reported that those labels stopped appearing, so that specific diagnostic signal became unavailable.
  3. What to do now: Verify OAI-SearchBot access, improve the opening of priority pages, then track Citation Share across a defined prompt set rather than relying on one screenshot.

Who does this affect most?

This affects small sites that depend on clear subject expertise rather than broad domain recognition. That includes niche B2B SaaS companies, local service firms, specialist publishers and ecommerce brands with pages that answer specific buying questions.

A small site can be a good source when it owns a narrow question. A tax software company may have a useful page for nonprofit bookkeeping. A regional HVAC firm may answer a local service question. A specialist retailer may publish a comparison with visible specifications. None needs to be a major publisher before its page can be useful to a search system.

The same result should not be read as a guarantee for small sites. A tight AI answer has limited room for sources. If several pages address the same topic, the system may cite the source that best fits its response. A page that is vague, technically blocked or slow to expose its main information gives the system less to work with.

The study also suggests that interface conditions matter. BotRank reported that free-account results leaned toward the observed in-house index, while paid thinking-mode results in its sample leaned much more toward Google scraping. A site may therefore appear in one context and not another. That is not necessarily a contradiction. It can be a different retrieval route.

This is why teams should avoid treating one result as a verdict on visibility. A prompt, account type, mode, location and time can all change the answer surface. Citation Share is more useful because it measures the percentage of relevant AI answers in a category that cite you, rather than treating a single answer as the full market.

  1. B2B SaaS teams should test category, alternative and comparison prompts that prospective buyers use.
  2. Local businesses should test service and location prompts that map to calls or bookings.
  3. Specialist publishers should protect clear explanations, original reporting and visible source material.
  4. Ecommerce teams should make product specifications and comparisons available in the initial page content.
Dated before-and-after interpretation of the observed ChatGPT Search change
PeriodWhat the evidence supportsWhat a small site should do
Before July 21, 2026Resoneo's observed pipeline labels, as reported by BotRank, suggested distinct in-house-index and Google-scraping retrieval paths.Use the available path data as a diagnostic, but do not treat it as an official ranking specification.
After July 21, 2026BotRank reported that the observed pipeline labels stopped appearing, reducing visibility into that specific retrieval signal.Use repeatable answer testing, cited-domain analysis and crawl checks instead of depending on hidden pipeline labels.
Ongoing documented requirementOpenAI says OAI-SearchBot is used to surface sites in ChatGPT Search, and opted-out sites will not appear in ChatGPT Search answers.Allow OAI-SearchBot for content intended to appear in ChatGPT Search, subject to your site's policy.

Why is crawl access the first thing to check?

Crawl access is the first thing to check because a strong page cannot appear in ChatGPT Search answers if its site opts out of the crawler OpenAI identifies for that purpose. OpenAI documents OAI-SearchBot separately from GPTBot, which means allowing SearchBot does not require allowing the training crawler.

This separation matters for teams with cautious content policies. OpenAI says webmasters can allow OAI-SearchBot for ChatGPT Search while disallowing GPTBot for foundation-model training. The policy decision and the search-visibility decision are related operationally, but they are not the same permission.

Start with robots.txt, then check the server configuration around it. The rule needs to be accessible to the crawler, and the site must not block the published OpenAI IP ranges or return an unhelpful response to the page itself. OpenAI says systems can take about 24 hours to adjust after a robots.txt update, so a change should not be judged immediately.

Google's crawler documentation is relevant when BotRank's reported Google-scraping path is part of the observed picture. Google says its common crawlers respect robots.txt for automatic crawls. A page that is unavailable or poorly served to a crawler can weaken visibility wherever a retrieval route depends on that crawl.

Technical access is not a citation strategy by itself. It is the baseline. Once it is sound, the editorial question becomes whether the page answers a query more directly than the alternatives available to ChatGPT.

  1. Check that robots.txt allows OAI-SearchBot on pages you want eligible for ChatGPT Search.
  2. Keep GPTBot policy separate from OAI-SearchBot policy, because they have different documented purposes.
  3. Confirm that important page content is available in the HTML response and is not hidden behind an avoidable interaction.
  4. Retest after the adjustment period OpenAI documents for robots.txt changes.

How should a small site structure a page for retrieval?

A small site should structure priority pages so the title, H1 and opening sentences answer the query without delay. BotRank reported that the observed in-house index appeared to keep a title and a short snippet of roughly 200 characters, which makes the first visible explanation unusually consequential.

Put the concrete answer before brand language. A page about accounting software for nonprofits should begin by saying what it covers, who it is for and the basis for the comparison. A page that starts with a slogan, a long company introduction or template labels gives a retrieval system less direct language to reuse.

Heading structure also matters because it tells both readers and machines what each part of the page addresses. Use a single descriptive H1, then question-shaped H2s that answer the questions a buyer or researcher is likely to ask. Do not rely on a visually styled label that is not a meaningful heading.

Make the evidence easy to inspect. A comparison page should show the criteria, the sources and the scope of the comparison in the visible page content. A local page should make its service area and qualification clear. A definition should give the definition first, then explain its limits and use cases.

Freshness should be factual, not cosmetic. Update a page when the underlying product, process or data changes. Show a date when it helps a reader evaluate the information, but do not use a date as a substitute for revising stale copy. The objective is a page that remains accurate and immediately useful when retrieved.

  1. Write an H1 that names the question or category precisely.
  2. Use the first paragraph to give the direct answer and scope.
  3. Expose core facts, comparison criteria and citations in visible page content.
  4. Remove low-information template text that appears before the page's answer.
  5. Review priority pages whenever a product, policy or source changes.

How should teams measure whether the response worked?

Teams should measure the response with a repeatable prompt set, not a single favorable answer. Record the exact prompt, engine, mode, date and cited domains. Then compare how often the brand is cited across the question universe that matters to its category.

Start with a focused set of prompts that is discovery and evaluation. For a B2B product, include category questions, alternatives and comparison questions. For a local business, include service-plus-location questions. Keep the language stable long enough to see whether a page change coincides with a change in citations.

Separate Answer Presence from Citation Share. Answer Presence shows whether the brand appears across relevant questions. Citation Share shows the percentage of those answers that cite the brand. A brand can be mentioned without receiving the source credit that helps it own the answer.

Track competitor citations too. The most actionable finding is often not that your site was absent. It is that another domain supplied the explanation, comparison or source material that your page could have supplied more clearly. That directs the next editorial update toward a specific gap.

Do not promise an outcome from the observed change. The BotRank report is evidence that small sites may be eligible for an observed retrieval path, not evidence that any given page will be cited. The responsible response is to remove technical blocks, improve clarity and measure the result over time.

  1. Define a stable prompt set around real category, comparison, local or product questions.
  2. Run prompts across relevant ChatGPT modes and record the conditions used.
  3. Log cited domains and URLs, not only brand mentions.
  4. Calculate Citation Share for your brand and named competitors.
  5. Use recurring findings to prioritize page-level technical and editorial work.

What should a team do this week?

This week, audit the pages that should answer your most important ChatGPT questions. The objective is not to make every page longer. It is to make the highest-intent pages accessible, specific and easy to cite.

First, confirm the OAI-SearchBot rule and preserve your intended GPTBot policy. Second, select a limited set of category, comparison, local or product prompts. Third, examine the opening 200 characters of the page most relevant to each prompt. The page should state the answer before it explains the brand.

Then compare the page with the domains that ChatGPT currently cites. Look for a concrete difference: clearer scope, a direct table, a better-supported claim, more current details or a more precise answer. Fix the gap that is visible. Do not manufacture data or make a broad claim merely because a competitor does.

The news reaction here is cautiously positive for small sites. The evidence challenges the idea that publisher status alone decides inclusion in an observed ChatGPT Search path. It does not remove the need for strong pages, sound crawl controls or ongoing citation measurement. Rankings got you found. Citations get you chosen.

  1. Audit robots.txt for OAI-SearchBot access on priority content.
  2. Choose a repeatable prompt set tied to buyer or customer questions.
  3. Rewrite weak H1s and opening answers on the pages those prompts should retrieve.
  4. Document cited competitors and the source material they provide.
  5. Measure Citation Share after the changes, then repeat the review.

Key takeaways

  • Small sites are not automatically excluded from ChatGPT Search by the absence of a publisher deal, based on the dated observed evidence.
  • OpenAI documents OAI-SearchBot as the crawler used to surface websites in ChatGPT Search.
  • Allowing OAI-SearchBot and disallowing GPTBot are compatible policy choices because OpenAI documents different purposes for the bots.
  • The first page content should answer the query clearly because BotRank reported a short observed retrieval snippet.
  • One ChatGPT result is a weak measurement method because retrieval can vary by prompt and mode.
  • Citation Share is a stronger operating metric than a yes-or-no visibility check because it measures cited presence across relevant answers.

Omnicite Editorial. "ChatGPT Search Index: Can Small Sites Compete?" The Citation Report, Omnicite. https://omnicite.co/blog/can-small-sites-compete-in-chatgpt-s-search-inde/

Sources

Source: BotRank

Resoneo reviewed 1,249 ChatGPT answers in July 2026, and BotRank reported no measured serving difference between partner and non-partner pages in an observed in-house retrieval path. BotRank, 2026-08-14

Source: OpenAI Developers

OAI-SearchBot is used to surface websites in ChatGPT Search, and sites opted out of it will not be shown in ChatGPT Search answers. OpenAI Developers, 2026-08-14

Source: Google for Developers

Google's common crawlers respect robots.txt rules for automatic crawls. Google for Developers, 2026-08-14

Frequently asked questions

Can a small site appear in ChatGPT Search?

Yes. OpenAI says sites allowed for OAI-SearchBot can be surfaced in ChatGPT Search. BotRank's August 2026 report also found no measured serving difference between partner and non-partner pages in one observed in-house retrieval path, although that study is not an official OpenAI specification.

Do publisher deals determine ChatGPT Search visibility?

The available observed evidence does not support treating a publisher deal as the universal gate for small-site inclusion. It does not establish that deals have no effect on any other outcome, such as rights, freshness or citation frequency.

Which OpenAI bot should a site allow for ChatGPT Search?

OpenAI documents OAI-SearchBot as the crawler used to surface websites in ChatGPT Search. GPTBot has a separate documented purpose related to content that may be used for foundation-model training.

How quickly does a robots.txt update affect ChatGPT Search?

OpenAI says it can take about 24 hours for its systems to adjust after a robots.txt update. Check the technical change first, then allow for that documented adjustment period before assessing it.

What should the top of a page say for ChatGPT Search?

It should state the page's answer and scope clearly. Use a descriptive H1 and an opening paragraph that directly addresses the query rather than leading with generic brand language.

How should a team track ChatGPT visibility?

Use a fixed set of relevant prompts, record the mode and date, log cited domains and calculate Citation Share. Compare results over time rather than treating one answer as a final result.