Back to Blog
AI VisibilityAugust 24, 202612 min read

Where Does ChatGPT Get Its Data? Training, Search and Sources

ChatGPT answers from three layers: training data with a cutoff, licensed publisher content, and live search via Bing and its own crawler.

Jess O'Malley, author at Sightivo

ChatGPT gets its information from three places: the data its model was trained on (with a fixed cutoff date), content OpenAI licenses from publishers and platforms, and live web search that retrieves pages through Bing's index and OpenAI's own crawler when a question needs current facts.

Key numbers (as of August 2026)

  • OpenAI's current flagship models, the GPT-5.6 family, have a knowledge cutoff of 16 February 2026; GPT-5.5, the ChatGPT Plus default, cuts off at December 2025; GPT-4o, still used on the free tier, at October 2023 (OpenAI model documentation; RankScope, verified August 2026).
  • ChatGPT search launched on 31 October 2024, is powered by Bing's index, and has been free to all users since 5 February 2025 (Search Engine Land).
  • OpenAI documents four crawlers: GPTBot (training), OAI-SearchBot (ChatGPT search), ChatGPT-User (fetches a user asked for) and OAI-AdsBot (ad landing pages) (OpenAI).
  • Publisher deals: Associated Press (July 2023), Axel Springer (December 2023), Financial Times (April 2024), News Corp (May 2024, reported at up to $250 million over five years) and Reddit (May 2024, reported at roughly $70 million a year) (Digiday timeline; Search Engine Land on Reddit).
  • Reddit is 16.7% of all ChatGPT citations, Wikipedia 8.9%, Forbes 3.3% (Ahrefs, July 2026).
  • 72.4% of pages ChatGPT cites open with a short "answer capsule", and about 91% of those capsules contain no links, in an audit of 15 domains (Search Engine Land, November 2025).
  • Whether a brand gets named correlates with branded web mentions at 0.664, versus "very weak" correlations for backlinks, across 75,000 brands (Ahrefs, December 2025).
  • ChatGPT had 900 million weekly users in February 2026 (TechCrunch).

Layer 1: training data

The model behind ChatGPT is trained on a large corpus that OpenAI describes as "publicly available information on the internet, information from third parties, and information from users" (OpenAI, how models are developed). In practice that means web crawls (OpenAI's own GPTBot plus datasets such as Common Crawl for earlier models), books and reference material, code, and licensed content.

Training happens once per model, so everything the model "knows" from this layer is frozen at a knowledge cutoff. The current cutoffs:

ModelKnowledge cutoffWhere you meet it
GPT-5.6 Sol / Terra / Luna16 February 2026Current API flagships; ChatGPT for newer tiers
GPT-5.5December 2025ChatGPT Plus default (August 2026)
GPT-5 (original)September 2024Earlier 2025–26 default
GPT-4oOctober 2023Free tier default

Table dated August 2026; check OpenAI's model documentation for the current list.

Two things follow. First, anything that happened after the cutoff — a new product, a rebrand, a price change — is invisible to this layer. Second, this layer is where ChatGPT's opinions about brands mostly live: if it names three CRMs for startups without searching, it is recalling which names appeared often, in which contexts, across the training corpus.

Layer 2: licensed and partner content

Since 2023 OpenAI has signed content deals that feed both training and live answers: the Associated Press (July 2023), Axel Springer's titles including Politico and Business Insider (December 2023), the Financial Times (April 2024), News Corp's Wall Street Journal, Times and others (May 2024), plus Vox Media, The Atlantic, Condé Nast and more. Reddit gave OpenAI real-time API access to its posts in May 2024.

For a brand, the practical effect is that coverage on partner publications and discussions on Reddit are both training data and retrieval candidates — one reason Reddit dominates ChatGPT's citations.

Layer 3: live web search

When ChatGPT decides a question needs current information — or the user clicks the search icon — it retrieves web pages and answers with citations. OpenAI's documentation says ChatGPT "may share disassociated search queries with the Bing search engine to return web results" and combines that with its own technology; Microsoft confirmed at launch that Bing powers ChatGPT search.

OpenAI's own crawlers sit on top of that:

CrawlerWhat it doesrobots.txt
OAI-SearchBotSurfaces sites in ChatGPT search resultsRespected; block it and you can't be cited from live search
GPTBotCollects public content for training future modelsRespected; blocking it opts you out of training
ChatGPT-UserFetches a specific page when a user's request triggers itUser-initiated; OpenAI says robots.txt rules "may not apply"
OAI-AdsBotChecks ad landing pagesRespected

OpenAI notes that if a site allows both GPTBot and OAI-SearchBot, "we may use the results from just one crawl for both use cases", and that robots.txt changes take about 24 hours to register (OpenAI crawler docs).

So the retrieval pipeline for a live answer is: query → Bing index (plus OpenAI's search index) → candidate pages → the model reads them → answer with inline citations. A page that Bing has never indexed cannot enter that pipeline, however well it ranks on Google.

How ChatGPT chooses which sources to cite

Two studies describe what actually gets cited.

By domain. Ahrefs' monthly-refreshed tracking of a broad set of US queries puts Reddit first at 16.7% of citations, Wikipedia at 8.9%, then Forbes (3.3%), Merriam-Webster (3.1%), Consumer Reports (2.8%), Healthline (2.7%) and Walmart (2.6%). Community discussion, reference sites, big editorial brands and retail dominate; individual company blogs are rare.

By page shape. The Search Engine Land audit of 15 domains found 72.4% of ChatGPT-cited posts open with a short direct answer (an "answer capsule" of roughly 120–150 characters) and about 91% of those capsules contain no links. Evertune's analysis of around 40,000 URLs found the median cited page runs about 940 words with four H2s and ten images, and that listicles account for roughly half of citations.

By brand. Whether ChatGPT names you — as opposed to citing a page — tracks how often your brand is mentioned across the web (0.664 correlation in Ahrefs' 75,000-brand study), with YouTube mentions slightly stronger and backlinks barely registering.

Common misconceptions

  • "ChatGPT reads the whole internet live." It doesn't. Most answers come from training data; live search runs only when the model judges it necessary or the user forces it, and then it reads a handful of retrieved pages.
  • "If I rank on Google, ChatGPT will cite me." Live retrieval runs on Bing's index. Google rankings help indirectly (Bing and Google largely agree on strong pages), but Bing indexing is the actual gate.
  • "Blocking GPTBot hides my site from ChatGPT." It opts you out of training. Live search uses OAI-SearchBot, a separate crawler with its own robots.txt rule; ChatGPT-User fetches may not honour robots.txt at all.
  • "ChatGPT knows my latest pricing." Only via live search or a partner feed. From training data it knows whatever was on the web before the cutoff — December 2025 for the Plus default, October 2023 on the free tier.

What it means for your brand

  1. Get indexed in Bing. Bing Webmaster Tools and IndexNow are table stakes for ChatGPT search.
  2. Allow the right crawlers. Let OAI-SearchBot through; decide separately about GPTBot.
  3. Write answer-first pages. A one- or two-sentence answer under each question heading is the passage that gets lifted.
  4. Be mentioned where ChatGPT already looks. Reddit threads, "best X" lists, review sites and major publications are the citation pool; a mention there is worth more than a backlink from a site ChatGPT never cites.
  5. Keep facts consistent and current. Cutoffs mean stale descriptions persist for a year or more; make the live web say the right thing so retrieval corrects the memory.

How to see which sources ChatGPT is using

You can check all three layers yourself in a few minutes.

  • Is it searching or recalling? Ask a category question ("best invoicing tools for freelancers") and look for the sources panel or inline citations. No citations means the answer came from training data; that is the layer where a brand's reputation is baked in. Re-ask with "search the web for this" to force retrieval and compare the two lists — they often differ.
  • Which pages did it read? With citations showing, open each source. You will usually find two or three listicles or review pages, a Reddit thread, and occasionally a vendor page. Those are the pages that decide the answer for that prompt; note the ones you are not on.
  • Is your site reachable? Check your robots.txt for GPTBot and OAI-SearchBot rules, and search your own domain in Bing. If Bing has not indexed a page, ChatGPT search cannot cite it.
  • Does it know your current facts? Ask "what does [your brand] do and what does it cost?" with search off, then on. The first answer shows what the model memorised before its cutoff; the second shows what the live web says. If they disagree, the live web is what you can fix.

Repeat the category prompts weekly, three times each, and log which brands and sources appear. That log is the baseline for everything else.

Timeline: how ChatGPT's data sources have changed

DateChange
Nov 2022ChatGPT launches, answering only from training data (cutoff September 2021)
Jul 2023Associated Press licensing deal — the first with a major news organisation
Aug 2023GPTBot documented, with a robots.txt opt-out for training
Dec 2023Axel Springer deal (Politico, Business Insider, Bild)
Apr–May 2024Financial Times and News Corp deals; Reddit real-time API access (May)
Oct 2024ChatGPT search launches on Bing's index for Plus and Team users; OAI-SearchBot documented
Feb 2025ChatGPT search free for all users in supported regions
2025–26GPT-5 family: cutoffs move from September 2024 (GPT-5) to December 2025 (GPT-5.5) and 16 February 2026 (GPT-5.6)

Each step added a way for current information to reach an answer without waiting for the next model — which is why a brand's live web presence now matters more than what the model happened to memorise.

Where Sightivo fits

Sightivo runs your buyer prompts on ChatGPT, Claude and Google AI Overviews, records which brands and which sources each answer used, and turns the gap into a prioritised list of pages to get mentioned on — with the people behind them and drafted outreach. The free AI visibility checker shows in a minute whether ChatGPT names you for five prompts in your category.

Common Questions

Does ChatGPT use Google or Bing?

Bing. OpenAI's documentation says ChatGPT may share disassociated search queries with Bing to return web results, and Microsoft confirmed Bing powers ChatGPT search. OpenAI supplements this with its own crawler (OAI-SearchBot) and licensed partner content. ChatGPT does not use Google's index.

Does ChatGPT browse the internet?

When needed, yes. Since October 2024 ChatGPT decides per question whether to search the web, and users can force a search. Outside those cases it answers from training data frozen at the model's knowledge cutoff.

What is ChatGPT's knowledge cutoff?

It depends on the model: 16 February 2026 for the GPT-5.6 family, December 2025 for GPT-5.5 (the Plus default in August 2026), September 2024 for the original GPT-5, and October 2023 for GPT-4o on the free tier. Live search fills in anything newer.

Does ChatGPT use Reddit?

Heavily. Reddit is the single most-cited domain in ChatGPT answers (16.7% of citations, per Ahrefs' July 2026 data), and OpenAI has had real-time API access to Reddit content under a licensing deal since May 2024.

Can I stop ChatGPT using my website?

Partly. Disallowing GPTBot in robots.txt opts your site out of training future models; disallowing OAI-SearchBot removes you from ChatGPT search results. ChatGPT-User fetches triggered by a user's request may not honour robots.txt. Blocking both crawlers also means ChatGPT can no longer cite you.

How does ChatGPT choose which sites to cite?

For live answers it retrieves candidate pages through Bing and its own index, then the model selects passages that answer the question directly — cited pages disproportionately open with a short answer capsule, are list- or comparison-shaped, and come from community, reference, editorial and retail domains. For brand recommendations without search, it recalls names that appeared often across the training corpus, which is why web mentions predict visibility better than backlinks.

Key Takeaways

  • Three layers: frozen training data (cutoff Feb 2026 for GPT-5.6, Dec 2025 for the Plus default, Oct 2023 on free), licensed content (AP, Axel Springer, FT, News Corp, Reddit and more), and live search on Bing's index plus OpenAI's crawler.
  • Four crawlers with separate rules: OAI-SearchBot for search, GPTBot for training, ChatGPT-User for user-triggered fetches, OAI-AdsBot for ads.
  • Reddit takes 16.7% of citations; cited pages open with a direct answer 72.4% of the time.
  • Being named tracks web mentions (0.664), not backlinks.
  • Be indexed in Bing, allow OAI-SearchBot, answer first, and get mentioned on the sources ChatGPT already cites.

Run the free AI visibility checker to see whether ChatGPT names you today, then use the generative engine optimization guide to work through the fixes.

Topics covered

where does ChatGPT get its datawhere does AI get its informationwhat is ChatGPT trained ondoes ChatGPT use BingChatGPT sourcesChatGPT knowledge cutoffOAI-SearchBotGPTBot

Written by

Jess O'Malley, author at Sightivo
Founder

Product leader who's launched 8 B2B SaaS products over the past 6 years. Experienced in taking products from 0 to 1 and scaling them. Built Sightivo out of frustration while doing backlink outreach for another startup—spent hours juggling spreadsheets and tools just to send a few emails. Decided to build something better and share it with others facing the same pain.

Ready to grow your search & AI visibility?

Sign up for Sightivo to track rankings, backlinks, and how often AI engines cite your brand.

Start free