[The Engines]
Should Your Brand Block AI Crawlers?
Cloudflare starts blocking AI training and agent crawlers by default on September 15, 2026, splitting bot traffic into Search, Agent and Training lanes. Block the wrong lane and you can lose the citations you were trying to earn.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
Cloudflare will block 'mixed-use' AI crawlers, the bots that blend search, training and agent duties, by default on any ad-bearing page starting September 15, 2026, splitting AI traffic into three enforced lanes: Search, Agent and Training. If you run ads and haven't set custom crawler rules, check them now: blocking the wrong lane can silence you in the ChatGPT and Perplexity answers you'd rather be cited in.
What changed with AI crawlers on September 15, 2026?
Cloudflare starts blocking Training and Agent category AI crawlers by default on every page that carries advertising, unless the site owner says otherwise. The cutover lands September 15, 2026, and it forces every AI crawler into one of three declared lanes: Search, Agent, or Training (Help Net Security, 2026-07-02).
The logic is simple. A Search or Answer bot fetches a page to cite it in a live response, the kind of visit that can turn into a citation in ChatGPT, Perplexity, or an AI Overview. A Training bot fetches the same page to feed a model's next training run, with no promise of a citation ever coming back. Agent bots sit in between, browsing on a user's behalf. Cloudflare's old rules treated all three the same. The new default does not, and any crawler that blends more than one purpose without declaring it gets treated as blocked too.
Cloudflare also replaced its Pay Per Crawl pilot with a Pay Per Use program, which compensates publishers when their content actually surfaces inside an AI answer rather than simply when a bot visits (TechCrunch, 2026-07-01). The timing is not an accident. Automated traffic, including AI crawlers, passed human traffic on Cloudflare's own network this year, reaching 57.4% of requests against 42.6% from people (The Motley Fool, 2026-08-07).
Who does the change affect?
It hits anyone running an ad-supported site through Cloudflare, plus every brand trying to build Citation Share across the engines that matter. New Cloudflare customers, new sites added by existing customers, and every site still on the free plan inherit the new default automatically. Paid customers who already set custom crawler rules keep them, but anyone who never touched the settings is about to have a decision made for them.
B2B SaaS teams and local service businesses both sit in the blast radius, just for different reasons. A SaaS blog wants to show up when ChatGPT is asked for the best tool in its category. A local business wants to surface when someone asks an engine for the best plumber in a given city. Both depend on Search and Answer crawlers reaching their pages, and both can lose that access by accident if they treat 'AI crawler' as one undifferentiated threat.
The instinct to block is not irrational. When an AI summary appears above a Google result, users click through to a traditional result in only 8% of visits, down from 15% when no summary appears (Pew Research Center, 2025-07-22). Publishers watching referral traffic shrink have real reason to want fewer bots on their servers.
Blocking is also not a clean fix on its own. In the fourth quarter of 2025, 30% of AI bot scrapes ignored the robots.txt file entirely, taking the pages anyway (The Register, 2025-12-08). A default rule change helps, but it does not replace an actual policy for which bots you want and which ones you don't.
| Crawler category | Before September 15, 2026 | After September 15, 2026 (ad-bearing pages) |
|---|---|---|
| Search and Answer bots (OAI-SearchBot, PerplexityBot, Google-Extended) | Allowed by default | Still allowed by default |
| Training bots (GPTBot, CCBot, and similar) | Allowed unless the site owner blocked them | Blocked by default |
| Agent bots (task and browsing agents) | Allowed unless the site owner blocked them | Blocked by default |
| Mixed-use bots blending more than one purpose | Allowed unless the site owner blocked them | Blocked by default until the operator splits the traffic |
How should you respond?
Set policy by what each bot does, not by whether its name contains the letters 'AI'. That single distinction keeps you visible to the engines that cite you while still letting you say no to the ones that only train on you.
- Pull your current robots.txt and, if you're on Cloudflare, your AI Crawl Control settings. Identify every crawler by name, not by a blanket 'AI' label.
- Keep Search and Answer bots open, including OAI-SearchBot, PerplexityBot, and Google-Extended for AI Overviews, if citation share is part of your growth plan.
- Decide on Training bots like GPTBot and CCBot separately, on your own licensing or IP terms, without dragging your Search bots into the same rule.
- If you run ads on Cloudflare and haven't touched crawler settings, check them before September 15, 2026. After that date, the platform decides for you.
- Re-verify after the cutover. A mixed-use crawler you never split into Search and Training can lose you both categories in one move.
What does this mean for citation share?
Rankings got you found. Citations get you chosen, and a citation can only happen if the crawler reading your page for an answer is allowed through the door in the first place. That is the practical stakes behind a robots.txt line most teams never look at twice.
None of this is about gaming an engine into citing you. It is about not accidentally locking the door on the crawlers that would cite you anyway, while still holding a real position on the ones that only want your content for training. Get that split wrong once, on the wrong page, and you won't see the loss in a dashboard. You'll just quietly stop showing up in the answer.
Key takeaways
- Cloudflare blocks Training and Agent category AI crawlers by default on ad-bearing pages starting September 15, 2026.
- The change only auto-applies to new customers, new sites, and free-plan sites; existing paid configurations are untouched.
- Click-through to a traditional search result drops to 8% when an AI summary appears, versus 15% without one, which is why publishers want to block bots in the first place.
- Blocking is not airtight: 30% of AI bot scrapes in Q4 2025 ignored robots.txt outright.
- Blocking every 'AI' crawler by name can remove Search and Answer bots along with Training bots, cutting your citation share by accident.
- Audit your crawler rules by bot purpose before September 15, 2026, and re-check them after the cutover.
Omnicite Editorial. "AI Crawlers: Should You Block Them in 2026?" The Citation Report, Omnicite. https://omnicite.co/blog/should-your-brand-block-ai-crawlers/
Sources
Source: Help Net Security
Cloudflare blocks Training and Agent category AI crawlers by default on ad-bearing pages starting September 15, 2026 Help Net Security, 2026-07-02
Source: TechCrunch
Cloudflare's new policy separates crawlers into Search, Agent and Training categories and launches a Pay Per Use program TechCrunch, 2026-07-01
Source: The Motley Fool
Automated traffic reached 57.4% of requests on Cloudflare's network versus 42.6% from humans The Motley Fool, 2026-08-07
Source: Pew Research Center
Click-through to a traditional search result fell to 8% when an AI summary appeared in Google results, versus 15% without one Pew Research Center, 2025-07-22
Source: The Register
30 percent of AI bot scrapes in Q4 2025 bypassed explicit robots.txt permissions The Register, 2025-12-08
Frequently asked questions
What is the difference between a Training crawler and a Search crawler?
A Training crawler, such as GPTBot or CCBot, fetches your page to feed a model's training data with no promise of a citation later. A Search or Answer crawler, such as OAI-SearchBot or PerplexityBot, fetches your page to pull facts into a live answer, which is the visit that can turn into a citation.
Will blocking AI crawlers stop me from being cited in ChatGPT or Perplexity?
It can, if you block Search and Answer bots along with Training bots. Blocking only the Training category leaves the door open for the crawlers that actually generate citations in live answers.
Does this change apply to my site if I'm not on Cloudflare?
Not directly. The September 15, 2026 default only applies to sites on Cloudflare's network. It's still worth checking your own robots.txt or CDN settings, since other providers are moving toward similar AI crawler categories.
What happens if I do nothing before September 15, 2026?
If you're a new Cloudflare customer, a new site, or still on the free plan, Cloudflare will block Training and Agent crawlers on your ad-bearing pages by default. Existing paid customers with custom rules already in place keep those rules.
Can I block AI training bots without losing citations?
Yes. Blocking Training bots like GPTBot or CCBot while explicitly allowing Search and Answer bots like OAI-SearchBot, PerplexityBot, and Google-Extended lets you hold a position on training data without losing visibility in live AI answers.
What is Cloudflare's Pay Per Use program?
It's the replacement for Cloudflare's earlier Pay Per Crawl pilot. Instead of paying publishers per bot visit, it compensates them when their content actually surfaces inside an AI answer.