[The Engines]

How to Optimize for AI Crawlers: Training vs. Answering

Cloudflare just split AI bot access into search, training and agent categories, and a same-week ranking report shows citation and ranking have officially decoupled. Here is what changed and what to do about it.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

On September 15, 2026, Cloudflare retired its single 'Block AI Bots' switch and replaced it with three separate controls for search, training and AI agent crawlers, blocking training and agent bots by default on ad-supported sites while leaving answer bots untouched. The same day, Optimization Theory's 2026 ranking factors report found that only 17 to 38% of AI Overview citations now trace back to a top-10 organic result, down from roughly 76% in mid-2024. Sites still running a blanket bot block risk losing citations they never meant to touch.

What changed for AI crawlers on September 15, 2026?

Cloudflare retired its single 'Block AI Bots' toggle and replaced it with three independent controls: one for search crawlers, one for AI training crawlers, and one for AI agents acting on a user's behalf. The change is live now, not proposed.

Ad-supported sites get the stricter default: search crawlers stay allowed, while training and agent bots are blocked unless the site owner turns them back on. Sites without ads keep everything allowed until the owner changes it. Mixed-use crawlers such as Googlebot, Applebot and Bingbot get evaluated under both policies at once, so a training block can catch a crawler a site owner assumed was safe for search.

The timing lines up with a second shift. Optimization Theory's report, published the same week, found the gap between ranking and citation widening: only 17 to 38% of AI Overview citations now come from a page ranking in the organic top 10, down from roughly 76% in mid-2024. Two changes landing together turn bot-level access control into a ranking-adjacent decision, not a security afterthought handled once and forgotten.

Who does the training-versus-answering split affect?

Everyone on Cloudflare's free plan, everyone launching a new site there, and everyone who never touched their AI bot settings gets the new defaults automatically. That covers a wide slice of the web without a single click from the site owner.

For B2B SaaS and tech growth teams, the stakes are direct: a site that never explicitly allowed OAI-SearchBot or PerplexityBot could now sit outside the exact 'best [category] tool' answers that Citation Share is built to measure. For local and service businesses, the same accident quietly removes them from '[service] in [city]' answers, and it never shows up as a lost keyword ranking, only as a citation that stops appearing.

Publishers and ad-supported marketplaces face the sharpest edge, since ad-supported pages inherit the stricter default. Sites without ads keep the old open posture, which is its own risk if keeping training crawlers out was ever the goal in the first place.

Cloudflare's AI crawler defaults, before and after September 15, 2026
Crawler categoryBefore Sept 15, 2026After Sept 15, 2026What to do now
AI training crawlers (GPTBot, ClaudeBot, Google-Extended)Blocked only if the site owner had manually flipped the old blanket 'Block AI Bots' switchBlocked by default on every ad-supported site, including new signups and all free-plan sitesLeave disallowed if that is the intent; it does not affect citations either way
Search or answer crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot)Frequently caught in the same blanket block as training botsAllowed by default, separated from training bots for the first timeExplicitly confirm these stay allowed; blocking one removes a site from that engine's answers
Mixed-use crawlers (Googlebot, Applebot, Bingbot)Treated as ordinary search crawlersEvaluated under both the search and training policy, and blocked from training if a site disallows trainingUse each provider's documented opt-out, such as Google-Extended, instead of a shared crawler rule
User-triggered agents (ChatGPT-User, Perplexity-User)Not distinguished from other bot classes in most robots.txt filesBlocked by default on ad-supported pages alongside training botsAllow them if pages should be fetched live when someone asks a chatbot about a specific link

How should you respond to the change?

Audit robots.txt bot by bot instead of trusting a blanket rule, then check the defaults Cloudflare just applied actually match what the site wants.

  1. Pull the current robots.txt and check whether an old wildcard rule blocks OAI-SearchBot, Claude-SearchBot or PerplexityBot along with the training bots it was written for years ago.
  2. Confirm ad-supported pages are not accidentally blocking ChatGPT-User or Perplexity-User if the goal is to have pages fetched live when someone asks a direct question.
  3. Check that key pages render content server-side, since several AI crawlers execute little to no JavaScript, and content that only appears after client-side hydration is effectively invisible to them.
  4. Refresh cornerstone pages on a real cadence: pages updated within 30 days earn roughly 3.2 times more citations, per Optimization Theory's 2026 report.
  5. Track citation share by engine rather than organic rank alone, since the two now move independently of each other.

What is the actual difference between a training crawler and an answering crawler?

A training crawler collects content to build or refine a model. An answering crawler, sometimes called a search or retrieval crawler, fetches content specifically to generate or index a response shown to someone right now. Blocking the first costs a site nothing in citations. Blocking the second can remove that site from an answer entirely.

OpenAI runs three separate bots for exactly this reason: GPTBot for training, OAI-SearchBot for search indexing, and ChatGPT-User for fetching a page live when a person asks about it directly. Anthropic and Perplexity follow a similar pattern, most visibly ClaudeBot for training against Claude-SearchBot for retrieval, with PerplexityBot crawling purely to answer.

What happens when a site blocks the wrong bot?

The most common mistake is inherited, not chosen. A wildcard disallow rule written years ago to stop scraping now silently catches bots that did not exist when the rule was written. Optimization Theory's report puts training-focused crawler traffic at 51.8% of the total, mixed-purpose at 35.7%, and search-only at 9.3%, so a rule that blocks anything with 'bot' in the name is overwhelmingly likely to hit a training crawler, but it still catches enough search-only traffic to matter.

The cost of that mistake is rising. The share of AI bot requests met with a 403 Forbidden response climbed to 9.64% in July 2026, up from 5.67% a year earlier, according to TechnologyChecker.io's analysis of Cloudflare Radar traffic against more than 4,000 robots.txt files. Crawl-to-referral ratios tell the rest of the story: GPTBot returns one referral for every 251 pages it crawls, PerplexityBot for every 289, and ClaudeBot for every 1,917. A single misapplied disallow rule affects those ratios directly, and it never announces itself as a lost ranking.

Is this good news or bad news for citation share?

Neither, exactly. The split makes it possible to keep training crawlers out without losing citations for the first time, which is a genuine improvement over the all-or-nothing switch it replaced. It also raises the cost of getting the configuration wrong, because the defaults now do the blocking automatically instead of waiting for a site owner to flip a switch.

There is no page two in an AI answer, and there is no partial credit for a page an answering bot was never allowed to fetch. For a team already tracking citation share across ChatGPT, Perplexity, Gemini and AI Overviews, the fix is administrative: match the allow list to intent, not to sentiment about AI training. For everyone else, September 15 is a reasonable day to find out what robots.txt has actually been doing since it was last edited.

Pages crawled per referral sent back, by AI bot (July 2026)
0958.51917251pages crawled per referral289pages crawled per referral1917pages crawled per referralGPTBot (OpenAI)PerplexityBotClaudeBot (Anthropic)

Source: TechnologyChecker.io robots.txt and Cloudflare Radar analysis, 2026-09-03

Key takeaways

  • Cloudflare replaced its single AI bot switch with three separate controls for search, training and AI agent crawlers on September 15, 2026.
  • Ad-supported sites now block training and agent bots by default while leaving search and answer bots allowed.
  • Mixed-use crawlers like Googlebot are evaluated under both policies, so a training block can catch a crawler assumed to be safe.
  • A wildcard robots.txt rule written for old scraping bots is now the most common way to accidentally block an answering crawler.
  • AI bot 403 rates rose from 5.67% to 9.64% year over year, and crawl-to-referral ratios vary by more than 7x between bots.
  • Blocking a training crawler costs nothing in citations; blocking a search or agent crawler can remove a page from AI answers entirely.

Omnicite Editorial. "AI Crawlers 2026: Training vs Answering Bots" The Citation Report, Omnicite. https://omnicite.co/blog/how-to-optimize-for-ai-crawlers-training-vs-answ/

Sources

Source: Cloudflare

Cloudflare replaced its single AI bot toggle with separate controls for search, training and AI agents, blocking training and agent bots by default on ad-supported sites Cloudflare, 2026-09-15

Source: TechnologyChecker.io

The share of AI bot requests met with a 403 Forbidden response rose from 5.67% in July 2025 to 9.64% in July 2026, with GPTBot, PerplexityBot and ClaudeBot showing crawl-to-referral ratios of 251:1, 289:1 and 1,917:1 TechnologyChecker.io, 2026-09-03

Source: Optimization Theory

Only 17 to 38% of AI Overview citations now come from a top-10 organic result, down from roughly 76% in mid-2024, and pages refreshed within 30 days earn about 3.2 times more citations Optimization Theory, 2026-09-15

Frequently asked questions

What is the difference between a training crawler and an answering crawler?

A training crawler, like GPTBot or ClaudeBot, collects content to build or refine a model and has no direct effect on what a chatbot cites today. An answering crawler, like OAI-SearchBot, Claude-SearchBot or PerplexityBot, fetches content specifically to generate or index the response a person is shown right now.

Did Cloudflare block all AI bots on September 15, 2026?

No. Cloudflare blocked training and AI agent bots by default only on ad-supported sites. Search and answering bots remain allowed by default on both ad-supported and non-ad sites.

Does blocking GPTBot hurt ChatGPT citations?

No. GPTBot only feeds OpenAI's model training. The bot responsible for ChatGPT's live citations is OAI-SearchBot, and ChatGPT-User fetches a page when a user asks about that specific link.

How do I check if my robots.txt is accidentally blocking answer bots?

Compare every disallow rule against the current bot list rather than trusting a rule written years ago, since names like OAI-SearchBot and Claude-SearchBot did not exist when many blanket rules were set.

Why does content need to render server-side for AI crawlers?

Many AI crawlers execute little to no JavaScript, so content that only appears after client-side hydration is invisible to them and cannot be cited, regardless of how the page looks to a person in a browser.

What is Citation Share?

Citation Share is the percentage of relevant AI answers in a category that cite a given business, tracked across ChatGPT, Perplexity, Gemini and Google AI Overviews.