[The Engines]
Understanding Cloudflare's New Crawler Policies and Their Impact on AI Citations
Cloudflare's AI crawler policies now distinguish Search, Agent and Training traffic. The change gives publishers a more precise way to protect content without blindly cutting off discovery.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
Cloudflare's new AI crawler policies replace a blunt block-or-allow decision with controls for Search, Agent and Training traffic. For brands that want AI citations, the practical default is to allow Search while making an explicit decision on Agent and Training access. The September 15, 2026 default change matters because blocking Training can also block mixed-purpose crawlers, including Googlebot, Applebot and BingBot.
What changed in Cloudflare's AI crawler policies?
Cloudflare changed AI crawler policies on July 1, 2026 by separating automated traffic into Search, Agent and Training categories. Previously, its managed Block AI Bots option focused on a broad blocking choice. The new model lets site owners decide based on what a crawler is doing with their content, rather than treating every AI-related request as identical. Cloudflare's announcement describes the controls as available to all customers, including the Free tier.
Search crawlers collect or index content so they can answer questions later. Agent traffic acts in real time for a person, such as a chat fetch bot or a browser-use agent. Training crawlers collect content to train or fine-tune a model. Those purposes overlap in the real world, which is why this is not simply a prettier bot dashboard. It changes the access decision a publisher can make.
For an editorial team focused on AI visibility, Search is the category closest to citation opportunity. A search system needs to discover, retrieve and understand a page before it can plausibly surface that page as a cited source. That does not create a promise of citations. It does mean a blanket block is no longer the only available policy.
Cloudflare also says its classification can assign multiple purposes to one crawler. This is the operational detail that deserves the most attention. A crawler that does both Search and Training is not safely treated as Search-only when a site blocks Training. The policy evaluates the crawler's relevant behaviors, not the publisher's preferred interpretation of the visit.
- Before July 1, 2026: a managed Block AI Bots option offered a much broader control over AI crawler traffic.
- After July 1, 2026: customers can manage Search, Agent and Training behavior separately.
- From September 15, 2026: new-domain defaults apply different treatment to those categories on pages that display ads.
What will change on September 15, 2026?
On September 15, 2026, Cloudflare says new domains onboarding to its service will block Training and Agent crawlers by default on pages that display ads, while Search crawlers remain allowed by default. That is a default for new domains, not a universal instruction for every existing site. Cloudflare's policy documentation states that customers can opt out before the change takes effect.
The policy reflects a clear distinction: ad-supported pages are built for human attention, while Search traffic can return visitors to the publisher. Cloudflare's configuration does not decide whether a given publisher should value citations, referrals, model training, agent access or some combination. It gives that publisher a more specific lever.
The catch is mixed-purpose crawlers. Cloudflare says crawlers that combine Search and Training will be blocked by configurations that block AI training, including the legacy Block AI Bots option. Its July announcement names Googlebot, Applebot and BingBot as examples of multi-purpose crawlers affected by the most restrictive applicable rule.
That means a content team cannot assume that selecting block Training leaves all search discovery intact. It must inspect the crawler classifications and test the results for the pages that matter. For a company trying to earn Citation Share, accidental loss of search access can reduce the pool of systems able to discover current content.
- New domains with ads: Training and Agent traffic will be blocked by default from September 15, 2026.
- New domains with ads: Search traffic will remain allowed by default.
- Mixed-purpose crawlers: a Training block can also block crawlers that serve Search and Training roles.
| Period | Policy approach | What it means for AI citations | What to do |
|---|---|---|---|
| Before July 1, 2026 | Broad managed Block AI Bots option | A broad block can reduce the ability of affected systems to discover content. | Audit any legacy broad block and identify affected crawlers. |
| From July 1, 2026 | Separate Search, Agent and Training controls | Search access can be evaluated apart from training and real-time agent use. | Set an explicit policy for each behavior. |
| From September 15, 2026 | New ad-carrying domains block Agent and Training by default, Search stays allowed | Training blocks can affect mixed-purpose Search and Training crawlers. | Opt out where needed, test crawler classification and monitor access. |
Who does the new policy affect most?
The policy affects every Cloudflare customer that publishes content and cares about how automated systems reach it. The immediate impact is greatest for ad-supported publishers, B2B companies with large documentation libraries, marketplaces, local businesses with location pages and teams that have previously enabled a broad AI-bot block without reviewing its consequences.
Publishers whose business depends on human referrals may reasonably want stronger controls over Agent and Training access. They should still separate that commercial choice from their discovery policy. A page that is unavailable to a relevant search crawler has a weaker path to appearing in an answer that cites sources. The correct decision depends on the site's business model, content rights and the specific crawler involved.
B2B SaaS and technology growth teams should care because their buyer questions often lead to comparison pages, integration documentation, security pages and implementation guides. These are the assets a search system needs to retrieve when someone asks which product fits a requirement. Blocking a mixed-purpose crawler without testing can make that material harder to find at the moment a buyer is evaluating options.
Local and multi-location businesses face a related risk. Service pages, location pages and structured contact information need to remain accessible to the systems that build answers for location-based questions. A crawler policy should support the business's visibility goals, not create an unnoticed technical barrier between an accurate page and a potential citation.
- Ad-supported publishers need to review the September default before it applies to a new domain.
- B2B teams need to protect discovery of technical, comparison and product information.
- Local businesses need to preserve access to pages that answer service and location questions.
- Any site using legacy Block AI Bots needs to check how mixed-purpose crawlers are treated.
How do these policies affect AI citations?
The policies affect the conditions for AI citations, not the citation result itself. Allowing Search crawlers gives relevant systems an access path to collect or index a page. It does not guarantee that ChatGPT, Perplexity, Gemini, Copilot or Google AI Overviews will cite that page. Citation selection still depends on the quality, coverage and freshness of the content available to each system.
That distinction matters because crawler controls can be mistaken for a citation tactic. They are not a way to manipulate models. They are a publishing-access policy. The useful question is whether a setting matches the business's desired balance between discoverability, real-time agent access and model training.
A practical citation strategy begins with pages that answer a clear question, show their evidence and stay current. A crawler policy then needs to avoid stopping legitimate search discovery. Omnicite calls the measurable outcome Citation Share: the percentage of relevant AI answers in a category that cite you. Access alone is not Citation Share, but inaccessible content has less opportunity to become source material.
Cloudflare's change makes a more measured posture possible. A publisher can allow Search while reviewing Training and Agent traffic separately. That is more useful than assuming every automated visit creates the same value or the same risk. It also makes configuration review part of editorial operations, especially after a domain migration, security change or bot-management update.
- Allowing Search supports discovery but does not promise an AI citation.
- Blocking Training is a rights and commercial decision that can affect mixed-purpose crawlers.
- Agent access should be evaluated separately from search discovery.
- Citation performance should be measured across the relevant answer surfaces, not inferred from one crawler setting.
What should you do before the new defaults take effect?
You should audit the current Cloudflare bot configuration, identify business-critical crawler access and record a deliberate policy before September 15, 2026. Start with the pages that carry the most commercial or editorial weight: product pages, comparison pages, documentation, research, service pages and location pages. This work is not glamorous, but it is cheaper than discovering after a traffic or citation drop that a configuration change caused it.
Next, list the crawler categories your site intends to allow or block. Do not use generic labels such as good bots or bad bots as the final policy. Use Cloudflare's Search, Agent and Training definitions, then document what each choice means for the site. If the site has advertising, include the new-domain default in that review. If the site does not have advertising, do not assume the ad-page default decides the correct configuration for you.
Then test the policy. Confirm how Cloudflare classifies the named crawlers that matter to your audience, particularly any crawler with mixed purposes. Cloudflare's Block AI Bots documentation says mixed Search and Training crawlers are affected by Training-block configurations. A correct policy on paper can still be wrong if the affected crawler is not the one the team assumed it was.
Finally, monitor outcomes after the change. Review crawl logs, search performance, referral patterns and the questions where your brand seeks citations. If a material shift appears, compare the date with policy changes before rewriting content. Publishing better pages is important. So is confirming that the pages are available to the systems expected to find them.
- Audit existing Cloudflare AI and bot settings before September 15, 2026.
- Classify desired access by Search, Agent and Training behavior.
- Check whether any named crawler has mixed purposes.
- Test access to important content routes after changes.
- Track crawl activity and citation visibility alongside content updates.
Should you allow Search crawlers and block Training crawlers?
For many publishers seeking AI citations, allowing Search while separately deciding on Training is the most defensible starting position. It keeps a path open for discovery while giving the publisher a direct policy choice about model training. It is a starting position, not a universal setting, because Cloudflare's mixed-purpose rule can make that combination block crawlers the team expected to retain.
The right answer depends on the exact crawlers, the site's commercial model and the value of its content. A publisher with subscription content may reach a different conclusion from a local service business that wants broad discovery. A software company with public documentation may need a different posture for documentation and for its logged-in application. Policy should follow the content and the outcome sought.
The key is to make the choice explicitly. Do not leave a legacy broad block in place because it once sounded prudent. Do not allow every crawler because it sounds open. The new controls exist precisely because Search, Agent and Training traffic have different effects on a publisher's content, audience and operating model.
- Preferred starting position: allow Search where citation discovery is a goal, then make separate decisions for Agent and Training access.
- Verify mixed-purpose crawler treatment before relying on that policy.
- Review the setting after major content, monetization or infrastructure changes.
What is the dated before-and-after of Cloudflare's change?
The before-and-after is straightforward: before July 1, 2026, Cloudflare's managed approach centered on a broad Block AI Bots option. After July 1, customers can manage AI traffic according to Search, Agent and Training behavior. The change becomes operationally sharper on September 15, 2026, when new domains with ads receive defaults that block Training and Agent traffic while allowing Search.
What to do about it is equally direct: do not treat the new settings as a one-click citation switch. Create a documented access policy, preserve intended search discovery, verify mixed-purpose crawler behavior and watch the results. That is how a crawler policy supports Citation Engineering without claiming it controls an answer engine's editorial choices.
- Before: broader Block AI Bots control for AI crawler traffic.
- After July 1, 2026: separate Search, Agent and Training controls for all Cloudflare customers.
- After September 15, 2026: new ad-carrying domains block Training and Agent by default, while Search remains allowed.
- Action: review policy, test mixed-purpose crawlers and monitor citation-related outcomes.
Key takeaways
- Cloudflare separated AI crawler management into Search, Agent and Training on July 1, 2026.
- The new controls are available to Cloudflare customers, including the Free tier.
- From September 15, 2026, new domains with ads will block Agent and Training traffic by default while allowing Search.
- A Training block can also block mixed-purpose crawlers that perform Search and Training roles.
- Crawler controls influence access and discovery, not a guaranteed AI citation outcome.
- Teams seeking citations should preserve intended Search access, test policy behavior and monitor results.
Omnicite Editorial. "AI Crawler Policies: Cloudflare's New Rules" The Citation Report, Omnicite. https://omnicite.co/blog/understanding-cloudflare-s-new-crawler-policies-/
Sources
Cloudflare announced separate Search, Agent and Training AI traffic controls for all customers, including the Free tier, on July 1, 2026. Cloudflare, 2026-07-01
Cloudflare documents September 15, 2026 defaults for new domains and explains that mixed Search and Training crawlers are affected by Training-block configurations. Cloudflare Developers, 2026-07-01
The supplied July 2026 industry update reported Cloudflare's category-level policy shift and highlighted the mixed-purpose crawler implication. Ghaith Blog, 2026-07-30
Frequently asked questions
What are Cloudflare's new AI crawler policies?
Cloudflare now lets customers manage AI traffic according to Search, Agent and Training behavior. The policy update was announced on July 1, 2026.
When do Cloudflare's new default settings begin?
Cloudflare says the new defaults begin on September 15, 2026 for new domains onboarding to Cloudflare. On pages that display ads, Agent and Training are blocked by default while Search remains allowed.
Will blocking Training crawlers affect search visibility?
It can. Cloudflare says mixed-purpose crawlers that combine Search and Training are blocked by Training-block configurations, including the legacy Block AI Bots option.
Do Cloudflare crawler settings guarantee AI citations?
No. The settings control access to a site. Allowing Search can support discovery, but no crawler setting guarantees that an answer engine will cite a page.
Should a publisher allow Search crawlers?
A publisher that wants content discovered for potential AI citations should evaluate allowing Search crawlers. It should also verify how the relevant crawlers are classified and whether they have mixed purposes.
What should existing Cloudflare customers do?
Review existing bot settings, document desired Search, Agent and Training access, test mixed-purpose crawler behavior and monitor important pages after any change.