ChatGPT's retrieval was measured pipeline by pipeline, and the source engine changes with the tier
A French agency captured the raw server streams behind ChatGPT answers while OpenAI was still tagging each search result with the name of the pipeline that produced it. They logged 5,245 snippets across 424 free-tier answers in instant mode, plus 16,407 search results from paid accounts in thinking mode, and cross-checked with instrumented test pages, server logs and the search API.
The split by question type, free tier, instant mode: stable information 99.9% from OpenAI's own index across 2,136 snippets. Business and place, 100% across 310 snippets, none from Google. Products 97.6% across 1,392. News is the exception at 52.6% own index and 46.1% scraped Google across 1,407 results. Ask the same questions on a paid account in thinking mode and it inverts: scraped Google carries 95-99%, and about 75% of the 16,407 thinking-mode results overall.
Three more numbers worth keeping. On business and place queries in instant mode, 25% of results reached the model with no snippet at all, just a title and a URL. The snippet, when there is one, caps at about 202 characters and is taken from the page body starting at the H1. And the model only actually opens a page for around 1.2% of the URLs it pulls, but when it does it cites that page 74% of the time, against 7% for pages it merely retrieved.
Judge it honestly. One agency, self-built instrumentation, not peer reviewed, and the corpus skews toward news questions. The authors flag that OpenAI stopped tagging results with the source pipeline around 21 July 2026, so anything after that is inferred from format signatures rather than read. Treat the shape as solid and the exact percentages as one measurement.