[The Engines]
How Cloudflare's New Crawler Policies Impact AI Citations
Cloudflare replaced one blunt AI bot switch with policy controls for Search, Agent and Training. That makes crawler decisions more precise, but also easier to get wrong.
The short answer
Cloudflare's new AI crawler policy lets site owners manage Search, Agent and Training traffic separately. For citation visibility, the safest starting point is to allow Search, review Agent access by use case and treat Training as a separate rights decision, while checking mixed-purpose crawlers before the scheduled September 15, 2026 defaults take effect.
What changed in Cloudflare's AI crawler policy?
Cloudflare replaced a broad AI-bot control with three behavior-based categories: Search, Agent and Training. The July 1, 2026 announcement says the controls are available to all Cloudflare customers, including those on the Free tier. Instead of treating every automated visitor as one kind of AI bot, Cloudflare now asks what that visitor is doing with the content.
Search covers crawlers that collect or index pages so they can answer questions later. Agent covers automation acting in real time for a person, such as a chat fetcher or browser-use agent. Training covers crawlers taking content to train or fine-tune a model. One crawler can carry more than one classification, which becomes important when a restrictive rule applies.
The before-and-after is bigger than a dashboard redesign. Before July 1, the prominent managed option focused on blocking AI bots associated with model training. After July 1, site owners can make separate policy choices based on purpose. That creates a practical split between discoverability, user-directed access and model development.
Cloudflare also changed the meaning of a Verified bot. Verified no longer means automatically allowed. It means the bot can be allowed when its relevant category is allowed. The category policy now decides access, so a trusted identity does not override a site owner's Search, Agent or Training choice.
Why does the new policy matter for AI citations?
The new policy matters because citation access and training access are no longer forced into the same decision. A publisher can preserve access for Search crawlers that index pages for later answers while taking a different position on Training crawlers. That is closer to how citation visibility actually works: an engine needs a route to discover, retrieve or reference a page before it can cite that page.
This does not mean allowing Search guarantees a citation. Cloudflare's control governs access at the site edge, not an AI engine's source selection. An allowed crawler can still ignore a page, fail to understand it or choose a stronger source. The policy removes or preserves a technical route. It does not award Citation Share.
The reverse risk is clearer. Blocking a crawler classified for Search can remove one route through which an answer engine finds and refreshes content. That can weaken eligibility for current citations, especially for new or recently updated pages. The exact effect depends on the crawler, its classifications and the engine's other discovery paths, so teams should test rather than claim a universal traffic or citation loss.
Agent access is different again. An Agent may fetch a page because a user asked it to complete a task now. Blocking that category could stop a live retrieval even when Search remains allowed. For product pages, documentation, booking flows and other pages an assistant may need to use on a person's behalf, Agent policy deserves its own review rather than inheriting the Training decision.
| Date | Policy state | Citation risk | What to do |
|---|---|---|---|
| Before July 1, 2026 | A prominent managed control focused on broadly blocking AI bots associated with training. | A broad decision offered limited control over distinct discovery, live-agent and training uses. | Record any legacy Block AI bots setting before migrating policy. |
| July 1, 2026 | Search, Agent and Training controls became available to all Cloudflare customers, including the Free tier. | A site can preserve Search access while treating Agent and Training separately, subject to mixed-purpose classification. | Allow Search by default, review Agent use cases and check every mixed-purpose crawler before blocking Training. |
| September 15, 2026 | For new domains, Training and Agent are scheduled to be blocked on ad pages while Search stays allowed. Restrictive rules will apply to mixed Search and Training crawlers. | A Training block can also catch named mixed-purpose crawlers and reduce discovery access. | Audit new-domain defaults, ad pages and mixed-purpose crawlers before the scheduled date, then test and monitor. |
Who does the Cloudflare change affect?
The change affects every Cloudflare customer deciding how automated systems may reach public content, but the risk is highest for publishers and businesses that depend on discovery or AI-assisted journeys. Existing customers can configure the new categories now. New domains face an additional scheduled default change on September 15, 2026, particularly on pages that display ads.
Editorial publishers must balance content rights with citation reach. B2B SaaS and technology teams should protect access to product pages, comparisons, documentation and research that answer engines may reference. Local and service businesses should pay attention to location, service and booking pages because an Agent may need those pages to answer or act for a user.
The policy also affects teams that previously enabled Cloudflare's legacy Block AI bots option and assumed traditional search crawlers were outside the decision. Cloudflare says mixed-purpose crawlers that combine Search and Training will be subject to the most restrictive applicable rule from September 15. The company names Googlebot, Applebot and BingBot as examples that can be blocked when Training is blocked.
That mixed-purpose rule is the sharp edge. A policy intended to stop model training can also restrict Search behavior when both purposes share a crawler identity. Site owners that rely on those crawlers should audit classifications and settings before changing a production rule.
What happens on September 15, 2026?
On September 15, 2026, Cloudflare plans to apply new defaults to new domains: Training and Agent crawlers will be blocked on pages that display ads, while Search crawlers will remain allowed. This is a scheduled policy change, not a claim that all existing sites will suddenly receive the same configuration.
The default reflects Cloudflare's view that ad-supported pages are designed for human attention, while Search is the behavior most likely to send visitors back. Whether that default fits a specific site is a business decision. A publisher may welcome Search discovery but still need Agent access for user-requested summaries, research or transactions. Another publisher may decide that live agent access does not return enough value.
Cloudflare's bot policy documentation says customers can opt out before September 15. The key review is not only the three category toggles. Teams must also inspect crawlers with more than one purpose because the most restrictive matching policy can win.
For citation teams, the deadline creates a clean audit point. Record the current policy, identify important crawlers, map each one to Search, Agent and Training, then approve the intended state. Do not wait for a fall in crawl activity or citation presence to reveal that a broad block caught a crawler you meant to allow.
Should you block Training but allow Search?
Allowing Search while blocking Training is the most defensible starting policy for many citation-focused sites, but it is not a universal setting. It expresses a clear intent: the content may be indexed and referenced for discovery, but should not be collected for model training. Cloudflare built the new taxonomy to make that distinction possible.
The catch is crawler overlap. If a crawler is classified as both Search and Training, a Training block may block the crawler entirely under Cloudflare's restrictive matching rule. That means the apparent policy and the effective network behavior can differ. Review the actual bot classification before assuming that a green Search control preserves access.
Agent access needs a page-level and journey-level decision. An open article may benefit when a user-directed agent retrieves it. An authenticated area, expensive endpoint or sensitive workflow may need tighter controls. The right policy can therefore combine broad Search access with selective Agent handling, rate limits and existing security rules.
Do not use Training access as a proxy for citation strategy. Training and current-answer retrieval are distinct uses in Cloudflare's taxonomy. A site can decline one without intending to disappear from answer engines. The operational goal is precise access: preserve the routes that support discovery and useful user actions, then restrict uses that do not match the publisher's terms.
How should you respond to the policy change?
Respond with an access audit before changing any crawler policy. Export or record the current Cloudflare settings, list the answer engines and conventional search engines that matter to the business, and identify which verified crawlers serve them. Then compare the desired policy with the effective result for every mixed-purpose bot.
First, keep Search allowed unless there is a documented reason to sacrifice discovery. Second, decide whether Agent visits help users complete useful actions on the site. Third, make the Training decision separately, based on content rights and business policy. Fourth, inspect ad-carrying pages and new-domain onboarding plans before September 15. Fifth, test representative URLs from outside the normal browser path.
Apply the change in stages. Start with a small set of pages or a monitored window where possible. Watch for changes in verified-bot requests, crawl errors, indexing freshness and answer presence. Keep a rollback record with the previous settings and the exact time of the change. That lets the team distinguish a policy effect from an unrelated shift in content or engine behavior.
Finally, align the technical setting with editorial work. Crawler access cannot rescue thin, stale or unsupported pages. Omnicite's citation approach remains quality, coverage and freshness at scale, not manipulating models. The policy should let legitimate discovery reach citation-grade content, while the content itself earns selection through clear answers, original evidence and maintained sources.
- Inventory current Search, Agent and Training settings.
- Identify mixed-purpose crawlers before blocking Training.
- Keep Search access open by default for citation discovery.
- Review Agent access against real user journeys and security limits.
- Record a baseline and a rollback point before production changes.
- Recheck settings before the scheduled September 15, 2026 defaults.
How can you tell whether the policy affected citations?
You can tell only by measuring access and citation outcomes before and after the policy change. Start with a dated baseline: crawler requests by category, successful responses, blocked responses, indexed-page freshness and the set of tracked prompts where the brand is cited. Keep the policy change time exact so later comparisons use the right boundary.
Separate leading indicators from outcomes. Crawler activity and response codes move first. Citation Count per day, Answer Presence and Citation Share may move later and can change for reasons unrelated to Cloudflare. A drop in crawler access is evidence that the rule changed retrieval. A drop in citations is an outcome that still needs comparison against content changes, competitor movement and engine volatility.
Use a controlled URL sample. Include newly published articles, recently updated pages, stable evergreen pages and pages used in high-value user journeys. Test whether important crawlers can fetch them, then monitor whether the pages remain present in answers. This is more useful than watching a single sitewide traffic number.
Do not overstate causation. If citations decline after a block, restore the prior policy for the affected crawler and watch whether access and answer presence recover. If they do not, investigate other causes. The point of the new policy is precision, and the measurement plan should be equally precise.
Key takeaways
- Cloudflare now separates Search, Agent and Training instead of forcing one broad AI-bot decision.
- Allowing Search preserves a discovery route, but it does not guarantee an AI citation.
- Blocking Training can also block mixed-purpose crawlers that serve Search.
- New domains are scheduled to receive different defaults on ad pages from September 15, 2026.
- Agent access should be reviewed against real user journeys, not treated as Training traffic.
- Measure crawler access and citation outcomes before and after every policy change.
Omnicite Editorial. "Cloudflare AI Crawler Policy and Citations" The Citation Report, Omnicite. https://omnicite.co/blog/how-cloudflare-s-new-crawler-policies-impact-ai-/
Sources
Cloudflare introduced separate Search, Agent and Training controls for all customer plans and announced new defaults for September 15, 2026. Cloudflare, 2026-07-01
Cloudflare documents the new defaults and the treatment of mixed-purpose crawlers under restrictive Training rules. Cloudflare Developers, 2026-07-01
The July 2026 AI SEO update identified Cloudflare's crawler taxonomy change as a material update for citation strategy. Ghaith Blog, 2026-07-30
Frequently asked questions
What are Cloudflare's three AI crawler categories?
They are Search, Agent and Training. Search indexes content for later answers, Agent acts in real time for a user, and Training collects content to train or fine-tune a model.
Does allowing Search guarantee AI citations?
No. It preserves a technical discovery route, but the answer engine still decides whether the page is useful, authoritative and relevant enough to cite.
Can blocking Training also block search crawlers?
Yes. Cloudflare says the most restrictive rule will apply to crawlers classified for both Search and Training, including named examples such as Googlebot, Applebot and BingBot.
What changes on September 15, 2026?
Cloudflare plans new defaults for new domains. Training and Agent crawlers will be blocked on pages displaying ads, while Search crawlers will remain allowed.
Should citation-focused sites allow Agent crawlers?
It depends on the user journey. Allow Agent access where live retrieval or action helps users, then protect sensitive, authenticated or costly routes with suitable security controls.
How should a site test a new crawler policy?
Record a baseline, change one policy dimension at a time, test representative URLs, monitor allowed and blocked requests, and compare Citation Count, Answer Presence and Citation Share after the change.