[The Engines]

Should You Block AI Crawlers? Pros and Cons for SEO

Bot traffic just overtook human traffic on the web for the first time, and publishers are rewriting their robots.txt files in response. Here is the dated timeline and the decision framework for what to block.

Explore this article with AI

Open a source-aware analysis with this article as the primary source.
ChatGPTClaudePerplexityGeminiGrokGoogle AI

The short answer

In June 2026, Cloudflare reported that bots, most of them AI crawlers and agents, made up 57.4% of web page requests, ahead of humans at 42.6%. The right response is not a blanket block: separate crawlers that send you citations from crawlers that only take your content, then set robots.txt rules by crawler type, not by a single 'AI' switch.

What changed with AI crawlers this year?

AI crawlers just crossed a line the open web was not built to handle. On June 4, 2026, Cloudflare's traffic data showed automated requests, most of them AI bots and agents, accounting for 57.4% of all web page requests, against 42.6% from humans (NBC News, citing Cloudflare, 2026-06-04). Cloudflare's own leadership had expected that crossover around 2027. It arrived about eighteen months early.

That milestone did not appear out of nowhere. Nine months earlier, on September 24, 2025, Cloudflare shipped its Content Signals Policy, a robots.txt extension that lets site owners state, machine-readably, whether a crawler may use their content for search indexing, live AI answers, or model training. The policy rolled out across more than 3.8 million domains on Cloudflare's managed setup, defaulting to allow search, block training, and stay neutral on live AI answers.

The pattern behind both events is the same. AI companies are sending more bots to fetch web content than ever, and site owners are responding by writing rules for the first time about who gets to take what. Progressive Robot's August 2026 reporting on the shift captured the mood directly: publishers are moving from allow-lists to block-lists, a reversal of thirty years of default-open behavior toward search engines.

Why is Wikipedia losing traffic, and is your site next?

Wikipedia is losing human traffic because fewer people click through once an AI system already answered their question. The Wikimedia Foundation reported that human pageviews fell roughly 8% for the March to August 2025 period compared with the same months in 2024, after the Foundation tightened its bot-detection to strip out non-human traffic disguised as real visits (Wikimedia Foundation, Diff blog, 2025-10-17).

The mechanism is measurable further upstream too. Pew Research Center analyzed a month of real browsing data from 900 U.S. adults and found that when a Google search produced an AI summary, people clicked a traditional result link only 8% of the time, versus 15% when no summary appeared. Clicks on a link inside the summary itself happened just 1% of the time (Pew Research Center, 2025-07-22).

Wikipedia is a useful case because it is heavily cited by AI answer engines and still lost traffic. If one of the most-cited sources on the internet is down 8%, sites with less citation share should assume the same pressure applies, only with less compensating visibility. This pressure reaches your site regardless of how well it is cited today. What matters is how you respond without cutting off the traffic and citations you can still keep.

AI crawler timeline: what changed and what to do about it
DateWhat happenedWhat to do
2025-09-24Cloudflare launches the Content Signals Policy, a robots.txt extension covering 3.8M+ domains, defaulting to allow search and block AI trainingCheck whether your host or CDN supports Content Signals, and set search, training, and AI answers separately
2025-10-17Wikimedia Foundation reports human pageviews down about 8% for March to August 2025 versus the same period in 2024Track referral and direct traffic trends monthly so a decline shows up before it compounds
2026-06-04Cloudflare data shows bots at 57.4% of web page requests, passing humans for the first timePull your own server logs and identify which crawlers are actually hitting your site before writing any block rule
2026-08-31Reporting shows major publishers shifting from allow-lists to block-lists for AI crawlersDecide crawler by crawler (search, training, agent), not with one blanket AI rule

Who does this actually affect?

It affects anyone whose growth plan assumes a steady stream of organic clicks. B2B SaaS and tech growth teams feel it when a prospect asks ChatGPT for the 'best [category] tool' and gets a named answer without ever visiting a site. The visibility that matters there is not a search ranking. It is whether the team's content shows up, gets cited, and gets named against competitors on that exact prompt.

Local, multi-location, and service businesses feel a parallel version of the same problem. Someone asks an AI assistant for the best plumber, dentist, or accountant in their city, and the assistant answers from whatever content it has already crawled and trusts. A business that never showed up in that answer set loses the call before the person ever opens a browser tab.

Publishers feel it hardest and first, since referral traffic is the business model. Reuters, Time, and People Inc. have already shifted toward block-lists rather than allow-lists for AI crawlers, per Progressive Robot's reporting. Every one of these groups needs a different answer to the same question: which crawlers are worth letting in, and which ones are pure cost.

Which AI crawlers should you block, and which should you let in?

Not all AI crawlers do the same job, so a single 'block AI' rule throws away the crawlers that send you citations along with the ones that only take your content for training. The useful split is by function: search crawlers that power live citations, training crawlers that ingest content for model weights, and agent crawlers that act on a user's behalf inside a session.

Treat each category on its own terms before you touch robots.txt. A crawler that can put your page in front of an AI answer today is worth keeping open. A crawler that only feeds a future model with no citation path back to you is the one worth restricting.

  1. Search and answer crawlers (OAI-SearchBot, PerplexityBot, Google-Extended in AI Overviews mode): these fetch pages to generate or ground live answers and can produce a citation back to your site. Generally allow.
  2. Training crawlers (GPTBot, ClaudeBot, CCBot, Bytespider): these ingest content to train or fine-tune models, with no direct citation event tied to the fetch. This is where the Content Signals Policy's 'block training' default applies for most publishers.
  3. Agent crawlers (Claude-User, ChatGPT-User): these act inside a live user session, often to complete a task the visitor initiated, closer to a proxy for a real person than a bulk scraper. Evaluate case by case rather than blocking outright.
  4. Legacy search bots (Googlebot, Bingbot): these still power traditional indexing and should stay open regardless of your AI crawler stance, since blocking them removes you from search entirely, not just from AI answers.

How should you actually respond, step by step?

Start by reading your server logs, not your assumptions. Most site owners can name GPTBot and Googlebot and stop there, while a dozen other crawlers are actively fetching pages with no plan behind who is allowed to.

Second, write crawler-specific rules instead of one AI blanket rule. Cloudflare's Content Signals Policy is the fastest path if you are already on Cloudflare: set search to allow, training to block, and leave AI answers neutral until you have evidence either way for your category.

Third, keep your highest-value pages accessible to the crawlers that generate citations, and keep them current. A stale page that an AI answer engine already indexed and quotes from is a liability, since outdated claims attributed to you stay live in AI answers long after you have corrected the source page.

Fourth, measure citation share, not just crawler hits. Blocking traffic is easy to log. Whether you are actually being named in AI answers across ChatGPT, Perplexity, Gemini, and AI Overviews for the prompts that matter to your category takes a different kind of tracking, and it is the number that should actually drive your robots.txt decisions.

Is blocking AI crawlers worth the tradeoff?

Blocking protects content from uncompensated training use, and it does nothing to grow the audience that is already choosing AI answers over search results. Those are two separate goals, and conflating them leads to the wrong rule.

There is no page two in an AI answer. If a category prompt gets asked a thousand times a month and your competitor is the one consistently cited, a strict block-everything stance on AI crawlers does not win that prompt back, it just removes you from contention entirely. The sites gaining ground right now are not the ones blocking hardest. They are the ones publishing enough well-sourced, current content that search and answer crawlers keep finding a reason to cite them, while training crawlers get turned away at the door.

The honest tradeoff is speed of decision versus quality of decision. A blanket block is fast and satisfying and wrong for most businesses that still want to be found. A crawler-by-crawler policy takes an afternoon of log review and gets updated as new bots show up, which is the real, ongoing cost of operating on a web where bots now outnumber people.

Key takeaways

  • Bots passed humans on the open web in June 2026, at 57.4% of page requests, arriving about eighteen months ahead of Cloudflare's own forecast.
  • Wikipedia's human pageviews fell roughly 8% between March and August 2025, a sign that even heavily-cited sources are not immune to AI answer traffic loss.
  • When a Google search shows an AI summary, click-through drops to 8%, versus 15% without one, per Pew Research.
  • The right response is not a blanket AI block. Separate search and answer crawlers, which can send citations, from training crawlers, which cannot.
  • Cloudflare's Content Signals Policy lets you set search, AI training, and AI answers as three separate rules instead of one.
  • Track citation share across ChatGPT, Perplexity, Gemini, and AI Overviews, not just crawler hit counts, since that is the number a robots.txt decision should actually be based on.

Omnicite Editorial. "Should You Block AI Crawlers? Pros and Cons" The Citation Report, Omnicite. https://omnicite.co/blog/should-you-block-ai-crawlers-pros-and-cons-for-s/

Sources

Source: NBC News (citing Cloudflare)

Bots made up 57.4% of web page requests in June 2026, passing humans for the first time NBC News (citing Cloudflare), 2026-06-04

Source: Cloudflare

Cloudflare launched the Content Signals Policy across 3.8 million domains, allowing granular control over search, AI training, and AI answers Cloudflare, 2025-09-24

Source: Wikimedia Foundation

Wikipedia human pageviews fell about 8% for March to August 2025 versus the same period in 2024 Wikimedia Foundation, 2025-10-17

Source: Pew Research Center

Click-through on traditional search links drops to 8% when an AI summary appears, versus 15% without one Pew Research Center, 2025-07-22

Source: Progressive Robot

Major publishers are shifting from allow-lists to block-lists for AI crawlers Progressive Robot, 2026-08-31

Frequently asked questions

What is an AI crawler?

An AI crawler is an automated bot run by an AI company that fetches web pages, either to train a model, to ground a live AI answer with current content, or to act inside a user's session as an agent. Examples include GPTBot, ClaudeBot, PerplexityBot, and Google-Extended.

Should I block GPTBot?

GPTBot is OpenAI's training crawler, not the one that powers live ChatGPT citations. Most publishers can block GPTBot for training without losing citation opportunities, since OAI-SearchBot handles the search and answer function separately.

Does blocking AI crawlers hurt my SEO?

Blocking traditional search crawlers like Googlebot or Bingbot hurts SEO directly, since it removes you from indexing entirely. Blocking AI training crawlers such as CCBot or GPTBot does not affect traditional search rankings, since those are separate bots with separate purposes.

What is the Cloudflare Content Signals Policy?

It is a robots.txt extension Cloudflare launched on September 24, 2025 that lets a site state separate permissions for search indexing, live AI answers, and AI model training, instead of one blanket allow or disallow rule for every bot.

How do I check which AI crawlers are visiting my site?

Pull your server or CDN access logs and filter by user agent string. Most AI crawlers identify themselves clearly, for example GPTBot, ClaudeBot, or PerplexityBot, which makes it straightforward to see actual traffic before writing any block rule.

Will blocking AI crawlers stop AI Overviews from summarizing my content?

Only if you block the specific crawler that feeds that feature, such as Google-Extended for Google's AI systems. Blocking one AI company's crawler has no effect on a different company's AI answers, since each vendor runs its own bots.