[The Engines]

How Will Cloudflare's AI Blocking Affect Your SEO Strategy?

Cloudflare's AI-training controls are no longer separate from search crawler access for Googlebot and Bingbot. Review your settings before September 15, 2026, or a past AI-blocking choice could interrupt indexing.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

Cloudflare will treat Googlebot and Bingbot as crawlers that collect content for both search indexing and AI training from September 15, 2026. If a Cloudflare site blocks AI training, those search bots can be blocked too. Audit your Cloudflare controls, robots.txt, and crawl evidence now so protecting content does not quietly remove it from search and reduce the pages available to earn citations.

What changed in Cloudflare's AI crawler access rules?

Cloudflare's change makes AI-training controls consequential for traditional search visibility. Search Engine Journal reports that, from September 15, 2026, Cloudflare will classify Googlebot and Bingbot as bots that both index for search and gather AI-training data. A site configured to block AI training can therefore block those search crawlers as well.

The practical risk is historical configuration. A team may have enabled Cloudflare's old Block AI Bots setting to restrict model-training crawlers, then moved on. Under the reported change, that old decision can affect Googlebot and Bingbot. The setting is no longer an isolated preference about training data. It becomes part of the technical path that determines whether major search engines can request pages.

This matters because crawl access comes before indexing, discovery, and citation opportunity. A strong page cannot be selected for a search result or cited in an AI answer if the relevant engine cannot reliably reach it. Rankings got you found. Citations get you chosen. Both depend on access remaining intentional.

Cloudflare had previously positioned its controls as a way to help site owners manage AI-training access. Its July 2025 announcement introduced managed robots.txt support and an option to block AI bots on ad-monetized areas. The new September treatment changes the operational question from which AI bots should be blocked to whether a control now catches search bots your business needs.

  1. Before September 15, 2026, teams could generally treat AI-training blocks as a policy distinct from Googlebot and Bingbot search access.
  2. From September 15, 2026, the reported Cloudflare classification can include Googlebot and Bingbot when a site blocks AI training.
  3. Teams should review Cloudflare bot settings and verify live crawl access before the effective date.

Who is most exposed to the change?

Sites behind Cloudflare that have enabled AI-training blocking are most exposed, especially where the owner assumes Googlebot remains allowed. The risk is not limited to publishers. B2B SaaS sites, multi-location businesses, ecommerce stores, documentation hubs, and lead-generation sites can all lose discoverability if critical pages stop being crawled.

Teams with several owners are also exposed. Marketing may manage SEO, security may manage Cloudflare, and legal may have asked for an AI-training restriction. None of those choices is unreasonable. The failure comes when the controls are reviewed in isolation and nobody confirms the resulting treatment of Googlebot and Bingbot.

The affected inventory is often broader than the homepage. Check product pages, location pages, comparison pages, help content, category pages, XML sitemaps, and new articles. A partial block can create a misleading picture: branded pages may still appear in search while newer or deeper URLs fail to enter the index.

AI visibility work has an additional dependency. Omnicite tracks Citation Share across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews, but no measurement layer can compensate for source pages that engines cannot access. AI crawler access is a publication and infrastructure decision, not only a bot-management setting.

  1. Cloudflare customers that previously enabled AI-training or AI-bot blocking should review their settings.
  2. Sites where Cloudflare, SEO, content, and legal decisions sit with different people need a shared access review.
  3. Businesses that rely on new pages, location pages, or programmatic content for discovery face greater exposure.
  4. Teams measuring citations should confirm crawler access before drawing conclusions from their results.
Cloudflare AI crawler access: before-and-after operating check
PeriodWhat the setting can meanSEO riskWhat to do
Before September 15, 2026AI-training and AI-bot blocks could be managed as separate controls from Googlebot and Bingbot search access.Teams may assume an old AI-blocking choice is isolated.Inventory every AI-related Cloudflare control and capture the approved policy.
From September 15, 2026Search Engine Journal reports Cloudflare will classify Googlebot and Bingbot as crawlers that index for search and gather AI-training data.A site that blocks AI training can block search crawlers and interrupt crawling or indexing.Confirm Googlebot and Bingbot treatment in Cloudflare, then test priority URLs and monitor search-console evidence.
After the changePolicy labels and real edge behavior can diverge from team assumptions.A silent bot block can reduce discovery and citation opportunity.Keep a quarterly access review, with an owner and retained test evidence.

Why can a training control affect SEO?

A training control can affect SEO because crawler labels no longer map cleanly to one business purpose. Cloudflare's Content Signals Policy distinguishes search, ai-input, and ai-train. Search covers building a search index and returning results with links and excerpts. AI input covers retrieval, grounding, and real-time use in generative answers. AI training covers training or fine-tuning models.

That separation is useful in principle. It lets a site express a preference such as search=yes and ai-train=no. Cloudflare's September change shows why the interface and enforcement logic still need close review. A bot can have multiple uses, while your strategy may need separate rules for indexing, answer retrieval, and training.

Google's crawler documentation says that common crawlers, including Googlebot, respect robots.txt rules for automatic crawls. It also warns that inappropriate HTTP responses can affect how a site appears in Google products. If Cloudflare blocks the request at the edge, the page may never reach the stage where Google can fetch, render, or reconsider it.

Do not treat robots.txt as a complete diagnostic. Cloudflare explains that robots.txt tells crawlers which content they may access, while enforcement controls can stop requests. Inspect the rendered robots.txt file, Cloudflare settings, firewall or bot events, server logs, and Search Console evidence together. One screen is not proof that AI crawler access is working.

  1. Search indexing is crawling that supports search results, links, and excerpts.
  2. AI input is content used for retrieval or grounding in generative answers.
  3. AI training is content used to train or fine-tune a model.
  4. Edge enforcement is a Cloudflare decision that can prevent a crawler request from reaching the origin.

What should your team do before September 15?

Start with an audit, then make the smallest safe change. Identify every Cloudflare zone that matters to organic search, record its AI-bot and AI-training settings, and name the person responsible for each setting. Do not switch a control off simply because it looks risky. First establish which content uses the business permits and which crawlers are required for search performance.

Next, test access on representative URLs. Include a homepage, a recently published article, a product or service page, a location page if applicable, and a page that has historically driven search traffic. Verify that Googlebot and Bingbot can receive the intended response through Cloudflare. Compare the results with Cloudflare event logs and origin logs so a cached browser response does not disguise a bot block.

Then inspect search health. In Google Search Console, review indexing status and recent crawl evidence for priority URLs. Watch for a shift in crawled pages, indexing failures, or a sudden gap between published URLs and indexed URLs. Bing Webmaster Tools can provide a corresponding check for Bing. Establish a baseline before the effective date so a later change is visible rather than debated.

Finally, separate policy from panic. You can choose to restrict particular AI crawlers without allowing every automated client. The goal is not maximum bot access. The goal is a deliberate policy that preserves the search and answer-engine paths your business needs, while documenting any limits on training use.

Complete the work in sequence: inventory Cloudflare zones, AI controls, robots.txt files, and their owners. Test Googlebot and Bingbot access on priority URLs. Compare Cloudflare events with origin logs and search-platform evidence. Record the approved policy, then repeat the checks after September 15, 2026.

How should you configure the before-and-after policy?

The right policy is explicit, testable, and connected to business intent. Cloudflare's Content Signals Policy provides a useful vocabulary: allow search where you need discovery, decide separately on AI input, and decide separately on AI training. The September change means Cloudflare dashboard settings must be validated against that intent, not assumed to match it.

For a site that depends on Google and Bing visibility, the default safe position is to preserve their search-crawling access unless a documented business decision says otherwise. If your organisation wants to limit model training, have the Cloudflare owner confirm that the configuration does not block Googlebot or Bingbot after the new classification takes effect. Keep screenshots or exported settings alongside the test results.

For pages that should not appear in search, use a purpose-built indexing decision rather than using a broad AI-blocking control as a proxy. That keeps the policy legible. It also avoids a common operational error: someone changes an AI setting for one reason and accidentally changes search visibility for another.

Treat this as a recurring control, not a one-time migration. Cloudflare, search engines, and answer engines can change crawler behavior, product labels, and support boundaries. Schedule a quarterly access review and rerun it whenever a bot policy, CDN configuration, robots.txt file, or site architecture changes.

  1. Sites that need Google and Bing discovery should preserve their indexing access while restricting AI training only where appropriate.
  2. Cloudflare owners should confirm actual bot treatment rather than relying on the label of a toggle.
  3. Teams should test a representative URL set before and after the policy date.
  4. Governance records should identify the owner, approved intent, test evidence, and review date.

What does this mean for citation strategy?

It means technical accessibility must sit underneath Citation Engineering. Content quality, coverage, and freshness give an AI system more reason to trust and cite a source. They do not help if the engine cannot reach the source in the first place.

For The Citation Report's audience, the immediate action is not to chase a loophole. It is to make the crawl policy match the visibility strategy. If you want category pages, comparison content, and original research to compete for citations, those pages need clean access for the engines you rely on. Then measure the result with Citation Share, Citation Count per day, Answer Presence, and Share of Voice.

There is no page two in an AI answer. There is also no reliable citation path through a page that an engine cannot fetch. Review the infrastructure before you judge the content program.

  1. Teams should preserve intended crawler access before measuring citation outcomes.
  2. Priority pages should remain fresh, crawlable, and internally connected.
  3. Citation Share should be tracked alongside crawl and indexing health.
  4. Any unexpected Googlebot or Bingbot block should be escalated as a visibility incident.

Key takeaways

  • Cloudflare's reported September 15, 2026 change can turn an AI-training block into a Googlebot and Bingbot access problem.
  • Teams should audit past AI-bot settings, not only controls changed this month.
  • Real bot access should be verified with Cloudflare events, origin logs, and search-console evidence.
  • Search indexing, AI input, and AI training need separate policy decisions.
  • Protecting content and preserving search visibility are compatible only when the configuration is explicit.
  • Citation Share depends on authoritative content being accessible to the engines that may cite it.

Omnicite Editorial. "AI Crawler Access After Cloudflare's Change" The Citation Report, Omnicite. https://omnicite.co/blog/how-will-cloudflare-s-ai-blocking-affect-your-se/

Sources

Source: Search Engine Journal

Cloudflare will classify Googlebot and Bingbot as crawlers that index for search and gather AI-training data from September 15, 2026, according to Search Engine Journal. Search Engine Journal, 2026-09-04

Source: Cloudflare

Cloudflare introduced managed robots.txt support and AI-bot blocking controls, and reported that about 37% of the top 10,000 domains had a robots.txt file. Cloudflare, 2025-07-01

Source: Cloudflare

Cloudflare's Content Signals Policy distinguishes search, ai-input, and ai-train as separate uses of content. Cloudflare, 2025-09-24

Source: Google Search Central

Google says common crawlers, including Googlebot, respect robots.txt rules for automatic crawls and that inappropriate HTTP responses can affect a site's appearance in Google products. Google Search Central, 2026-09-06

Frequently asked questions

Will Cloudflare block Googlebot if I block AI training?

Search Engine Journal reports that from September 15, 2026, Cloudflare will treat Googlebot and Bingbot as crawlers that both index for search and gather AI-training data. A site set to block AI training can therefore block them. Confirm the live treatment in your own Cloudflare account before the date.

Does this change affect Bing SEO too?

Yes. The reported change names Bingbot alongside Googlebot. Review Bing access, not only Google access, if your Cloudflare configuration blocks AI training or AI bots.

Can robots.txt prove that Googlebot can access my site?

No. Googlebot generally respects robots.txt for automatic crawls, but a Cloudflare edge rule can still block the request. Check robots.txt, Cloudflare events, origin logs, and search-platform crawl evidence.

Should I allow every AI crawler to protect SEO?

No. The goal is an intentional policy, not unrestricted automated access. Preserve the crawlers needed for your search and answer-engine strategy, then apply specific limits that the business has approved.

What pages should we test first?

Test a homepage, a recent article, a core product or service page, a priority location page, and a page with established search traffic. These give a practical view of whether the policy affects both new and established inventory.

How does crawler access relate to AI citations?

Crawl access is a prerequisite for discovery. Citation Engineering still depends on quality, coverage, and freshness, but content cannot compete for Citation Share when the engines you depend on cannot reach it.