[The Engines]
Understanding ChatGPT's Labrador Index for Better AI Citations
ChatGPT quietly started serving a growing share of citations from its own search index, Labrador, instead of leaning only on Bing. Here is what the change means for anyone tracking AI citations.
Explore this article with AI
Open a source-aware analysis with this article as the primary source.The short answer
ChatGPT runs its own retrieval system called Labrador, a family of indexes covering web pages, PDFs, YouTube, news, arXiv, Wikipedia, local listings, finance, legal, medical content, shopping and images. For roughly two months in 2026 a leaked result_source field confirmed Labrador sits alongside Bing and scraping partners, and a page ranking well in Google is no longer automatically a page ChatGPT has indexed.
What is ChatGPT's Labrador index?
Labrador is not one index. It is OpenAI's own family of vertical indexes, separate retrieval systems built for general web pages, PDFs, YouTube, news, arXiv papers, Wikipedia, local listings, finance, legal and medical content, shopping and images, each one storing full page content alongside crawl dates and publication dates.
The name surfaced almost by accident. For roughly two months in 2026, a result_source field appeared in ChatGPT's server-side events and carried one of four values: Labrador, Bright, Oxylabs or SERP. The first is OpenAI's own index. The other three point to the scraping vendors and search partners ChatGPT has leaned on for live retrieval, according to reporting from Peec AI published 2026-09-18.
That field was never meant to be public. It disappeared once the pattern got noticed, but not before researchers logged enough sessions to map which verticals leaned on which source. A comparison prompt about software tools pulled from a different mix of sources than a local search for a nearby service, which is the first hard evidence that ChatGPT treats categories of queries differently at the retrieval layer, not just at the answer layer.
- General web pages
- PDFs and academic sources including arXiv and Wikipedia
- News and local listings
- Finance, legal and medical content
- Shopping and product pages, plus images
What changed in ChatGPT's indexing this month?
What changed is that ChatGPT stopped depending only on Bing and its scraping partners for live answers. Reporting from Peec AI and The GEO Show, both published in September 2026, traces Labrador's build-out back to 2023, a timeline that surfaced through Google's antitrust testimony, and describes OpenAI actively A/B testing its own index against partner search results with experiments like 'prefer-index-over-serp-v3' running in shopping.
The crawl behind it is not small. Peec AI reports ChatGPT crawling at roughly 35,000 requests per hour as it builds out coverage across those vertical indexes. Loved by AI's roundup, published 2026-09-07 with an update on 2026-09-14, frames the practical shift bluntly: a page Google knows about is not automatically a page ChatGPT knows about.
None of this happened as a single announcement. It surfaced in pieces, through a leaked field, through court testimony about a rival's antitrust case, and through independent crawl monitoring, which is why most publishers only found out about Labrador well after OpenAI had already been running it for years. There is no page two in an AI answer, and a page Labrador never crawled cannot appear on page one either.
The table below lines up the two states, how ChatGPT's retrieval worked before Labrador was confirmed and what this month's reporting shows now.
| Signal | Before Labrador surfaced | After 2026 reporting confirmed it |
|---|---|---|
| Primary retrieval source | Bing and scraping partners such as Bright Data and Oxylabs supplied most live citations | A growing share of citations comes from Labrador, OpenAI's own index family (Peec AI, 2026-09-18) |
| Visibility into source routing | No public signal on which system supplied a given source | A result_source field named Labrador, Bright, Oxylabs and SERP for about two months in 2026 (Peec AI, 2026-09-18) |
| Vertical coverage | One general web crawl passed through search partners | Separate indexes for web, PDF, YouTube, news, arXiv, Wikipedia, local, finance, legal, medical, shopping and images (Peec AI, 2026-09-18) |
| Crawl independence | Crawl pace set by Bing and partner schedules | OpenAI runs its own crawler, reported at roughly 35,000 requests per hour (Peec AI, 2026-09-18) |
| Google to ChatGPT correlation | Ranking well in Google largely predicted a ChatGPT citation | A page Google indexes is not automatically a page Labrador has indexed (Loved by AI, 2026-09-07) |
Who does the Labrador index affect?
It affects anyone who assumed a Google ranking would carry over to an AI citation. That assumption held while ChatGPT leaned on Bing and partner search results. It holds less now that a second crawler, with its own schedule and its own vertical priorities, decides separately what gets stored and surfaced.
Two groups feel this fastest. B2B software and tech teams competing for 'best [category] tool' answers need content Labrador's web and comparison-relevant indexes actually crawl, not just content that ranks. Local, multi-location and service businesses need visibility in the local and shopping verticals Labrador tracks on their own terms, separate from a Google Business Profile or a Google Maps ranking.
The effect compounds for anyone publishing at low volume or on a slow refresh cycle. A single crawler missing a page is a rounding error. Two crawlers with different schedules, priorities and vertical logic missing the same page is a page that effectively does not exist for a meaningful share of AI answers. Coverage across every engine, not just the one with the biggest search share, is what separates a site with real answer presence from one that only looks visible in a single dashboard.
- Teams relying on Google rank as their only AI visibility signal
- Sites with strong backlinks but thin or stale pages Labrador has little reason to crawl deeply
- Local and service businesses assuming a Google listing automatically carries into ChatGPT's local index
- Publishers whose update cadence is too slow for a second crawler to bother revisiting often
How should you respond to Labrador?
Respond by treating ChatGPT's crawler as a first-class indexer, not an afterthought sitting behind Google. That means publishing and refreshing content on a schedule Labrador's crawl can actually catch, structuring pages for the specific vertical they belong to instead of one generic catch-all, and checking citation presence per engine rather than assuming Google visibility settles the question.
None of this requires guessing at how Labrador ranks anything. OpenAI has not published ranking criteria, and nobody outside the company can claim otherwise. What is verifiable is the split itself: separate indexes, a separate crawler, and a result-routing field that, for a couple of months this year, proved it.
The practical shift is measurement, not just publishing habits. Citation share, the percentage of relevant AI answers in a category that cite a given site, now has to be tracked across engines individually, since a strong showing on Google AI Overviews says nothing about whether Labrador has even crawled the page ChatGPT would need to cite. Answer presence across the full question universe matters more than a single ranking snapshot once retrieval itself has split into separate systems. Share of voice against named competitors is worth watching too, since a competitor whose pages Labrador crawls deeply can out-cite a site that only optimized for Google.
- Publish and refresh often enough that a second crawler finds something new when it returns
- Structure content for the vertical Labrador tracks separately, such as local, shopping or finance, instead of one generic page
- Track citation share across ChatGPT, Perplexity, Gemini and AI Overviews individually rather than as one undifferentiated bucket
- Treat a strong Google ranking as necessary, not sufficient, for AI citation
Key takeaways
- ChatGPT runs its own index family, internally called Labrador, alongside Bing and scraping partners.
- A result_source field briefly named Labrador, Bright, Oxylabs and SERP for about two months in 2026, confirming the split.
- Labrador covers web, PDF, YouTube, news, arXiv, Wikipedia, local, finance, legal, medical, shopping and image content as separate indexes.
- OpenAI's crawl behind Labrador runs at a reported 35,000 requests per hour and traces back to a 2023 start.
- A Google ranking no longer guarantees a ChatGPT citation, so content needs to be crawlable and current on its own terms.
- Publishers should check citation presence per engine instead of treating 'AI search' as one bucket driven by Google rank.
Omnicite Editorial. "ChatGPT Citations: What the Labrador Index Changes" The Citation Report, Omnicite. https://omnicite.co/blog/understanding-chatgpt-s-labrador-index-for-bette/
Sources
Source: Peec AI
ChatGPT exposed a result_source field naming Labrador, Bright, Oxylabs and SERP for about two months in 2026, and Labrador covers separate web, PDF, YouTube, news, arXiv, Wikipedia, local, finance, legal, medical, shopping and image indexes, with a crawl rate of roughly 35,000 requests per hour. Peec AI, 2026-09-18
Source: The GEO Show (GEOforge)
OpenAI is building proprietary crawl and index infrastructure for ChatGPT called Labrador, reducing reliance on Bing and making ChatGPT's own crawler a first-class indexer alongside Googlebot. The GEO Show (GEOforge), 2026-09-07
Source: Loved by AI
OpenAI's indexing effort traces back to 2023 per Google antitrust testimony, and a page Google indexes is not automatically a page ChatGPT's Labrador index has crawled. Loved by AI, 2026-09-07
Frequently asked questions
What is ChatGPT's Labrador index?
Labrador is OpenAI's own family of search indexes for ChatGPT, covering general web pages plus separate indexes for PDFs, YouTube, news, arXiv, Wikipedia, local listings, finance, legal, medical content, shopping and images.
How was the Labrador index discovered?
A result_source field briefly appeared in ChatGPT's server-side events for about two months in 2026, carrying one of four values: Labrador, Bright, Oxylabs or SERP, with Labrador identifying OpenAI's own index (Peec AI, 2026-09-18).
Does a good Google ranking still guarantee a ChatGPT citation?
No. Reporting confirms that a page Google indexes is not automatically a page Labrador has indexed or crawled, since the two systems now run on separate schedules and priorities (Loved by AI, 2026-09-07).
How fast is ChatGPT crawling the web through Labrador?
Peec AI reports a crawl rate of roughly 35,000 requests per hour as OpenAI builds out Labrador's vertical coverage (Peec AI, 2026-09-18).
When did OpenAI start building Labrador?
Reporting ties the project's origin to 2023, a timeline that surfaced through Google's antitrust testimony rather than an OpenAI announcement (Loved by AI, 2026-09-07).
What should site owners do differently now?
Publish and refresh content on a schedule a second crawler can catch, structure pages for the specific vertical Labrador tracks, and measure citation presence across ChatGPT, Perplexity, Gemini and AI Overviews separately rather than assuming Google rank covers all of it.