[The Engines]
How to Tailor Your SEO Strategy for Different AI Engines
OpenAI rewrote its ChatGPT crawler rules on December 9, 2025, and a single robots.txt policy no longer covers every AI engine. Here is what changed, who it affects, and how to respond engine by engine.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
OpenAI rewrote its ChatGPT crawler rules on December 9, 2025, and robots.txt no longer controls ChatGPT-User the way it used to. Every AI engine now runs on its own logic: ChatGPT rewards crawler access and clean extraction, Perplexity rewards breadth, and Google's AI Overviews still run on classic SEO fundamentals. Treat 'AI SEO' as one strategy per engine, not one strategy for all of them.
What changed in how AI engines read your site?
OpenAI rewrote the rules for its ChatGPT crawlers on December 9, 2025, and the change matters more than the quiet rollout suggested. The company updated its official bot documentation and removed a promise that had shaped robots.txt strategy since GPTBot launched in August 2023: that ChatGPT-User would respect a site's robots.txt file the same way GPTBot and OAI-SearchBot do.
Before the revision, OpenAI's documentation stated that a shared set of robots.txt tags applied to all three agents. After it, only OAI-SearchBot and GPTBot are bound by those tags. OpenAI's stated reason is that ChatGPT-User requests are triggered by a live person asking a question, so the company now treats that traffic as human browsing rather than autonomous crawling, meaning robots.txt rules may not apply at all. The same update stripped language tying OAI-SearchBot to training data collection, expanded ChatGPT-User's documented scope to cover Custom GPT requests and GPT Actions, and confirmed that OAI-SearchBot and GPTBot may now share the results of a single crawl instead of hitting a site twice.
None of this happened in isolation, and none of it will be the last change. Perplexity runs its own bot, PerplexityBot, alongside a separate Perplexity-User agent that generally ignores robots.txt outright. Google keeps Google-Extended, a training-only opt-out with no effect on AI Overviews eligibility, separate from the standard Googlebot that still drives ranking and inclusion. Three companies, three sets of rules, and no shared standard between them. A team that still writes one robots.txt policy and calls it 'AI SEO' is already behind.
The pace is the real story. OpenAI has touched its crawler documentation multiple times since GPTBot's 2023 launch, each revision quietly reshaping what a blocked or allowed bot actually means for visibility. Waiting for a stable rulebook is not a strategy. Checking the current one on a schedule is.
Who does this actually affect?
Anyone whose growth plan depends on being cited, recommended, or quoted by an AI engine is affected, whether or not they have noticed yet. If a chatbot has ever named a competitor instead of you in answer to a buyer's question, the crawler and rendering rules behind that answer already decided the outcome. There is no page two in an AI answer, so a technicality in a bot's documentation can cost a deal nobody saw slip away.
Two groups feel it fastest, for different reasons. B2B SaaS and tech growth teams lose deals quietly when ChatGPT recommends a competitor for 'best [category] tool' and the buyer never opens a second tab to check. Local, multi-location, and service businesses lose the same way when someone asks an AI engine for the best plumber, dentist, or agency in their city and a competitor's page answers first. Both groups need the same underlying proof, and neither gets it from a Google rank alone.
- B2B SaaS and tech growth teams: proof lives in citation share on category and comparison prompts, plus share of voice against named competitors.
- Local and service businesses: proof lives in citation share on '[service] in [city]' prompts, AI Overviews presence, and the calls or bookings that follow.
| Element | Before Dec 9, 2025 | After Dec 9, 2025 |
|---|---|---|
| ChatGPT-User and robots.txt | Documentation stated shared robots.txt tags applied to GPTBot, OAI-SearchBot, and ChatGPT-User | Only OAI-SearchBot and GPTBot follow robots.txt; ChatGPT-User is treated as user-initiated browsing that robots.txt rules may not cover |
| OAI-SearchBot description | Linked to appearing in ChatGPT results and to training data collection | Framed only as the bot that surfaces websites in ChatGPT's search features; training language removed |
| ChatGPT-User scope | Covered general user-initiated fetches | Explicitly extended to Custom GPT requests and GPT Actions |
| Crawl coordination | Not documented | OAI-SearchBot and GPTBot may share the results of one crawl across both use cases |
How should you respond, engine by engine?
Treat each engine as its own channel with its own mechanics. Blocking, allowing, or restructuring content for one engine does nothing for the others, and a single generic bot-blocking file will not fix any of them.
- ChatGPT: confirm OAI-SearchBot can crawl and render your pages, since Pepper's engine-by-engine analysis of ChatGPT citation behavior weights crawler access at roughly 35%, rendering at 25%, and extractability at 25%, with freshness and structured data covering the rest. Do not assume robots.txt alone controls ChatGPT-User traffic anymore. For finer control over that agent, look at authentication, rate limiting, or an edge tool such as Cloudflare's bot management instead.
- Perplexity: optimize for breadth, not depth. PerplexityBot indexes for search, and Perplexity-User frequently pulls content regardless of robots.txt, so the safer lever is making more of your relevant pages easy to parse rather than gating one flagship page and hoping it carries the whole category.
- Google, meaning Gemini and AI Overviews: the fundamentals still apply. Google-Extended only governs training use and has no bearing on whether a page can appear in an AI Overview, so standard technical SEO, crawlability, and structured data remain the actual levers here, not a training opt-out.
- Copilot: it leans on Bing's index, so classic on-page and technical SEO still carry weight. Do not assume a page that ranks in Bing automatically gets cited in a Copilot answer. The extraction step still has to work, the same way it does for ChatGPT.
Why do the rules keep moving?
Because the underlying systems still fail more often than most teams assume. A May 2026 study cited in Pepper's guide to engine-by-engine optimization found that retrieval failures, not model reasoning errors, caused more than 70% of chatbot mistakes across 2,100 test questions. When a crawler cannot reach a page, or a renderer cannot parse it, the model has nothing accurate to cite, no matter how strong the writing underneath is.
Coverage matters as much as access. Muck Rack's May 2026 analysis of more than 25 million links cited across ChatGPT, Claude, and Gemini found that earned media, meaning independent articles, reviews, and journalism that mention a brand, accounts for 84% of AI citations, while paid and advertorial content accounts for roughly 0.3%. Schema markup and a clean robots.txt get you into the room. Being written about by someone else outside your own site is still what gets you quoted once you're in it.
Put those two findings together and the priority order gets simple. Fix access and rendering first, because a crawler that cannot reach the page cannot cite anything on it. Earn outside coverage second, because that is the material an AI engine actually trusts. Structured data and a tidy FAQ block help at the margins, but they do not substitute for either one.
What should you do this week?
Start with the check that costs nothing and takes ten minutes: read your own robots.txt file line by line, and do not assume it says what it said a year ago.
- Pull your current robots.txt and confirm GPTBot, OAI-SearchBot, PerplexityBot, and Googlebot are not blocked by accident.
- Stop relying on robots.txt to manage ChatGPT-User traffic. Add authentication, rate limiting, or edge-level bot rules if that agent needs restricting.
- Check that your most important pages render without JavaScript-only content, since a crawler that cannot render a page cannot cite it.
- Audit where your brand already earns unpaid mentions, and prioritize outreach and PR over adding more schema markup.
- Track citation share by engine, not just search rank, so you can see which engine is actually citing you and which one is not.
Source: Pepper, 'How to optimize your website for ChatGPT and Perplexity, engine by engine', 2026-09-10
Key takeaways
- OpenAI changed its crawler documentation on December 9, 2025, and ChatGPT-User is no longer reliably bound by robots.txt.
- GPTBot controls training data use, OAI-SearchBot controls whether ChatGPT's search has can cite your pages, and the two settings are independent.
- Google-Extended only affects AI training, not inclusion in Google's AI Overviews.
- Retrieval failures, not model errors, caused more than 70% of chatbot mistakes in a 2026 study cited in Pepper's engine-by-engine guide, so crawl access and clean rendering come first.
- Earned media drives 84% of AI citations across ChatGPT, Claude, and Gemini, compared with roughly 0.3% for paid content, per Muck Rack's May 2026 research.
- Measure citation share separately for each engine instead of one blended visibility score, since a page can win on one engine and disappear on another for unrelated reasons.
Omnicite Editorial. "AI Engine Optimization: Tailor Your SEO by Platform" The Citation Report, Omnicite. https://omnicite.co/blog/how-to-tailor-your-seo-strategy-for-different-ai/
Sources
Source: PPC Land
OpenAI revised its ChatGPT crawler documentation on December 9, 2025, removing robots.txt compliance for ChatGPT-User PPC Land, 2025-12-09
Source: Muck Rack
Earned media accounts for 84% of AI citations across ChatGPT, Claude, and Gemini, versus roughly 0.3% for paid content Muck Rack, 2026-05-07
Source: Pepper
ChatGPT, Perplexity, and Google weigh citation factors differently, and retrieval failures caused over 70% of chatbot errors in a 2,100-question study Pepper, 2026-09-10
Frequently asked questions
What exactly did OpenAI change on December 9, 2025?
OpenAI revised its crawler documentation so that only OAI-SearchBot and GPTBot are bound by robots.txt. ChatGPT-User, the agent that fetches pages during a live chat session, is now described as user-initiated browsing that robots.txt rules may not cover.
Does blocking GPTBot stop my site from appearing in ChatGPT answers?
No. GPTBot governs training data collection. The bot that controls whether your pages can be cited in ChatGPT's search has is OAI-SearchBot, and blocking that one is what removes a site from those answers.
Does Google-Extended affect AI Overviews?
No. Google-Extended is a training-only opt-out. It does not change whether a page can be crawled, indexed, or surfaced in Google's AI Overviews, which still runs on standard Googlebot access and normal SEO signals.
Is one robots.txt policy enough for every AI engine?
No. OpenAI, Perplexity, and Google each define their bots differently and change the rules on their own schedule, so a policy written for one engine can leave another engine blocked or unrestricted by accident.
What matters more for AI citations, technical SEO or PR?
Both, but earned media carries more weight than most teams expect. Muck Rack's May 2026 research found earned media behind 84% of AI citations across ChatGPT, Claude, and Gemini, far ahead of paid content, so crawl access has to be paired with real outside coverage.
How do I measure whether these changes are working?
Track citation share for each engine separately, meaning the percentage of relevant AI answers on your category's prompts that cite you, rather than one combined score across ChatGPT, Perplexity, Gemini, and AI Overviews.