hello@afrazalam.com Kolkata, India
Follow on:
AI Marketing

How Does ChatGPT Choose Its Sources? A 3-Engine Breakdown

How ChatGPT, Perplexity, and Google AI Overviews actually pick the pages they cite, straight from each provider's own documentation, plus what to fix first on your own site.

Last updated: September 2026 by Afraz Alam, Founder of Afraz Alam Marketing Services. I run AI search visibility work for US businesses, and this is the question I get asked most: how does ChatGPT choose its sources, and why isn't it choosing mine?

How does ChatGPT choose its sources? In search mode it rewrites your question into targeted queries, sends them to partner search indexes (OpenAI names Microsoft and Shopify), ranks the results on relevance and reliability, and cites the pages that survive. Perplexity and Google AI Overviews do the same job with 3 different pipelines, which is why a page can be cited by one and invisible to the other 2.

A Shopify store owner posted this in the Shopify Community last year: "Why does ChatGPT cite Reddit and gift guides for my products instead of my actual Shopify pages?" That one sentence is the whole problem. The engines are not ignoring you. They are picking sources using rules you have never seen written down.

This guide lays those rules out, engine by engine, using each provider's own documentation. No guessing, no recycled LinkedIn theories.

The short answer: 3 engines, 3 source pipelines

ChatGPT, Perplexity, and Google AI Overviews all end in the same place, a paragraph with links under it. They get there 3 different ways.

ChatGPT search leans on partner search indexes and a ranking layer OpenAI keeps private. Perplexity runs its own live retrieval first and hands the results to a language model second. Google AI Overviews fan your query out into many related searches against Google's own index, then pick supporting pages from the results.

Different pipelines mean different winners. If you treat "AI search" as one thing and optimize once, you are optimizing for none of them. That is the mistake behind most answer engine optimization advice I read.

How ChatGPT, Perplexity, and Google AI Overviews choose sources: 3 different pipelines compared side by side
Three engines, three source pipelines. Same question, different candidate pools.

How does ChatGPT choose its sources?

ChatGPT has 2 different ways of answering, and they source information completely differently. Most confusion about ChatGPT sources comes from not knowing which mode you are looking at.

Two modes, two different answers

Mode 1 is the model answering from training data. There are no live sources here. When ChatGPT answers this way, it is recalling patterns from text it was trained on, which is why it can name a brand it read about in 2024 and miss a competitor that launched in 2026. Ask it for a source in this mode and it will often invent one. That is where "chatgpt fake sources" and "chatgpt hallucinating sources" come from, both real Google autocomplete phrases.

Mode 2 is ChatGPT search. This is the mode with real citations, and it is the only mode you can influence with your website. When search runs, ChatGPT goes out to the live web through partner indexes and cites what comes back.

What OpenAI actually says

OpenAI's ChatGPT search help article is short, but it gives away the mechanism. It says ChatGPT search "sometimes partners with other search providers" and "typically rewrites your query into one or more targeted queries that it sends those providers." Microsoft and Shopify are the 2 partners named.

On ranking, the exact line is: "ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed."

On eligibility, OpenAI says to "allow OAI-SearchBot to crawl the site" and confirm your host or CDN allows traffic from its published searchbot IP addresses. OAI-SearchBot is the crawler for search results and links. It is separate from GPTBot, which OpenAI uses for training. Blocking GPTBot to protect your content while leaving OAI-SearchBot open is a real, supported choice.

Read those 3 statements together and the picture is clear. If you are not indexed and ranking in Bing for the query, you are not in the candidate pool for ChatGPT search. If you are a Shopify merchant, your product data has a second door in through the Shopify partnership.

Why Reddit and listicles keep winning

Back to the Shopify owner's question. Reddit threads and "best gifts" listicles get cited for product queries because the retrieval step is looking for pages that answer the user's actual question. The user did not ask "tell me about Brand X." They asked "what is a good gift for a runner." A Reddit thread with 40 people comparing options answers that. A product page with 3 bullet points and a price does not.

The ranking layer also favors what OpenAI calls reliable sources. In practice that means pages with third-party consensus, recognizable domains, and content that reads like an answer rather than an ad.

What a small brand can influence

You cannot buy placement. You can control 4 things: being crawlable by OAI-SearchBot, ranking in Bing for the questions your buyers ask, having pages that answer those questions in plain language in the first 100 words, and getting mentioned on the third-party pages ChatGPT already trusts. I show the page-level rewrite in How to Get Cited by ChatGPT, with a real before and after.

How Perplexity picks what to cite

Perplexity is the most transparent of the 3 about the fact that it searches, and the least transparent about how it ranks.

Retrieval first, model second

Perplexity's own help center describes the order of operations: it "searches the internet, gathering information from authoritative sources like articles, websites, and journals," then uses language models (it names GPT-5 and Claude 4.6 Sonnet) to understand the query and write the answer. Every answer "includes numbered citations linking to the original sources."

Retrieval first means the model never sees your page unless retrieval found it. Page structure, freshness, and a clear match between your heading and the user's question decide whether you enter the pool. Writing quality decides whether you get quoted once you are in it.

PerplexityBot vs Perplexity-User

Perplexity runs 2 different fetchers, and its bot documentation spells out the difference. PerplexityBot is "designed to surface and link websites in search results on Perplexity." It respects robots.txt and Perplexity recommends allowing it. Perplexity-User "supports user actions within Perplexity." When a user asks a question, it "might visit a web page to help provide an accurate answer and include a link to the page," and the documentation states plainly that "this fetcher generally ignores robots.txt rules."

That second fetcher is why blocking Perplexity in robots.txt does not reliably keep you out of Perplexity answers. It is also why Cloudflare publicly reported in August 2025 that it had observed Perplexity using undeclared crawlers to get around no-crawl directives. Whatever you think of that behavior, it changes the strategy: you cannot opt out cleanly, so you are better off making sure the page it fetches is the right one.

PerplexityBot vs Perplexity-User: which Perplexity fetcher respects robots.txt and which one ignores it
Perplexity runs 2 fetchers. Only one of them respects robots.txt.

Where Perplexity differs from ChatGPT

3 practical differences. Perplexity cites more sources per answer, typically 5 to 10 numbered citations against ChatGPT's smaller set. Perplexity weights recency harder because it fetches live for every query, so a dated "Last updated" line and current-year facts matter more here. And Perplexity is less dependent on Bing rankings, because it runs its own retrieval, so a page that is buried on Bing still has a shot if it matches the question tightly.

Where does Google AI Overview get its information?

From Google's index. That is the whole answer, and Google says so in its own documentation. The interesting part is how it picks from that index.

Query fan-out

Google's Search Central page on AI features says both AI Overviews and AI Mode "may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources, to develop a response." It then says that while the response is being generated, "our advanced models identify more supporting web pages, allowing us to display a wider and more diverse set of helpful links associated with the response than with a classic web search."

Translate that: a user asks 1 question, Google silently runs 6 or 8 related searches, and the pages that rank for those sub-questions become the candidate pool. You can be cited in an AI Overview for a query you do not rank for, if you rank for one of the sub-questions Google fanned out to.

Why page 1 is necessary, not sufficient

Google states the only requirement: your page must be "indexed and eligible to be shown in Google Search with a snippet." So traditional SEO is the entry ticket. But the fan-out means the page that wins is the one that answers a specific sub-question cleanly, not the one that ranks number 1 for the broad head term. A 4,000-word guide that buries the answer in paragraph 12 loses to a 900-word page that answers it in the first line.

What Google says, and what it does not say

Google is explicit that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary," and that "there's also no special schema.org structured data that you need to add." Take that at face value. Schema will not get you in.

What Google does not say is how it chooses between 2 eligible pages. It does not publish the weighting. Anyone selling you a "guaranteed AI Overview placement" formula is selling you something Google itself says does not exist.

Side by side: how the 3 engines choose sources

Question ChatGPT search Perplexity Google AI Overviews
Where do candidate pages come from? Partner search indexes (Microsoft, Shopify named) Perplexity's own live retrieval Google's own index via query fan-out
What gets you into the pool? Crawlable by OAI-SearchBot, ranking in Bing Tight match to the question, fresh content Indexed and eligible for a Google snippet
Citations per answer Fewer, often 3 to 5 More, often 5 to 10 Varies, shown as link cards
How much does recency matter? Medium High Medium, depends on the query
Third-party vs first-party bias Strong pull toward Reddit, reviews, listicles Moderate, will cite brand pages that answer directly Follows what already ranks
Can robots.txt keep you out? Yes, block OAI-SearchBot Not reliably, Perplexity-User ignores it Yes, standard Google controls apply
What to fix first Bing indexing and answer-first pages Freshness dates and question-matched headings Rank for the sub-questions, not just the head term

What this means for a DTC brand, a coach, and a home service business

If you sell products online. Your product pages are losing to Reddit and gift guides because they do not answer buying questions. Add a real "who this is for and who it is not for" section, a comparison against the 2 alternatives your buyers actually consider, and a plain-English answer to the top 3 questions from your support inbox. If you are on Shopify, keep your product data clean, because that partnership is a second route into ChatGPT.

If you sell coaching or a course. Your offer page is a sales page, and none of the 3 engines cite sales pages. They cite pages that explain. Write the page that answers "how does [your method] work" in the first 80 words, with your name on it, and link it from the offer page. That is the page that gets cited. The offer page gets the click.

If you run an HVAC, plumbing, or electrical company. Google AI Overviews matter most to you, and they pull from what already ranks locally. Your Google Business Profile and your service pages are the candidate pool. The fan-out question is not "best plumber Phoenix," it is "why is my AC blowing warm air," and the local company whose page answers that is the one that gets cited.

What to fix first, in order

  1. Confirm all 3 crawlers can reach you. Check robots.txt and your CDN rules for OAI-SearchBot, PerplexityBot, and Googlebot. One blocked bot removes you from one engine entirely.
  2. Check you are indexed in Bing, not just Google. Bing Webmaster Tools takes 10 minutes to set up. Without Bing, ChatGPT search cannot find you.
  3. Rewrite your top 5 pages to answer in the first 100 words. Question as the heading, direct answer as the first paragraph, proof under it. Every engine's retrieval rewards this.
  4. Add a visible "Last updated" date and keep it honest. Perplexity in particular weights recency, and a dated page beats an undated one on the same topic.
  5. Rank for the sub-questions, not only the head term. Map the 6 to 8 questions Google would fan out to for your main keyword, and make sure one page on your site answers each.
  6. Get mentioned where the engines already look. A single real mention in a Reddit thread, an industry roundup, or a review site does more for ChatGPT citations than 10 new blog posts on your own domain.
Six fixes, in order, to get cited by ChatGPT, Perplexity, and Google AI Overviews
The 6 fixes that move a site into all 3 candidate pools, in the order I run them.

If you want to see where you stand before doing any of this, run the AI search visibility audit first. It takes 20 minutes and it tells you which of the 3 engines is the real problem.

What I am not sure about

None of the 3 companies publishes a ranking specification. OpenAI says "multiple factors." Perplexity says "authoritative sources." Google says its models "identify more supporting web pages." The citation counts and recency weights in the table above come from my own testing across client and personal sites, not from a published formula, and they will drift as the products change.

I also cannot tell you how the Shopify partnership ranks one merchant above another inside ChatGPT. OpenAI has not published that. If someone tells you they know, ask them for the document.

The verdict

So, how does ChatGPT choose its sources? Through partner search indexes, a private ranking layer, and a strong preference for pages that answer the question rather than pitch a product. Perplexity does its own retrieval and rewards freshness. Google AI Overviews fan out your query and pull from pages that already rank for the pieces.

3 engines. 3 pipelines. 1 thing they all reward: a page that states the answer first and proves it second. Build that page and you have a shot at all 3. Skip it and no amount of schema, llms.txt files, or "AI optimization" will save you.


Want to know which of the 3 engines is actually ignoring you?

I run SEO, AEO, and AI search visibility work for US businesses on a free 30-minute strategy call. You bring the site, I show you where it drops out of each pipeline. No pitch deck.

Book a free strategy call

Sources cited in this article: OpenAI, ChatGPT search; Perplexity, How does Perplexity work; Perplexity, Crawlers and bots; Google Search Central, AI features and your website. All checked 29 September 2026.

Frequently Asked Questions

In search mode, ChatGPT rewrites your question into targeted queries, sends them to partner search providers (OpenAI names Microsoft and Shopify), and ranks the results on what OpenAI calls "multiple factors intended to help users find relevant, reliable information." The pages that rank get cited. Your site must be crawlable by OAI-SearchBot to be eligible.

From 2 places, depending on the mode. Without search, it answers from training data and has no live sources. With search on, it pulls live pages through partner indexes and cites them. Only the second mode can be influenced by your website.

Both happen. In search mode the citations are real links to real pages. In plain chat mode, if you ask for a source, the model can generate a plausible-looking citation that does not exist. If there is no clickable link, treat the source as unverified.

From Google's own search index. Google says AI Overviews use a "query fan-out" technique, running multiple related searches and then identifying supporting pages from those results. The only requirement to be eligible is that your page is indexed and can show a snippet in Google Search.

Yes. AI Overviews show supporting links alongside the generated text, and Google states it displays "a wider and more diverse set of helpful links" than a classic result page. You can be linked for a query you do not rank for directly, if you rank for one of the related searches Google ran.

Not reliably, and Google does not claim they are. The summary is generated by a model from the pages it retrieved, so an error on a cited page or a misread of it becomes an error in the overview. Check the linked sources for anything that matters.

Perplexity searches the live web first, then passes the results to a language model that writes the answer with numbered citations. It uses 2 fetchers: PerplexityBot for search indexing, which respects robots.txt, and Perplexity-User, which fetches pages on demand when a user asks and generally ignores robots.txt.

Because the user asked a question, not for your brand, and the Reddit thread answers the question. ChatGPT's ranking favors relevant, reliable pages, and a thread where real people compare options reads as both. The fix is a page on your site that answers the same buying question directly, plus real mentions on the third-party pages ChatGPT already trusts.
Afraz Alam

Afraz Alam

Digital Marketing Consultant

Afraz Alam is the founder of Afraz Alam Marketing Services, a digital marketing consultant helping businesses grow through SEO, Google Ads, Meta Ads, AI-driven marketing, and conversion-focused website development. He works directly on every account himself, no account managers, no handoffs, and reports on leads, calls, and revenue, not likes or impressions. His approach stays data-driven and transparent: real strategy, real reporting, no guaranteed rankings or inflated promises.

Prefer done-for-you over do-it-yourself?

Reading is a great start — but if you'd rather have it handled, book a free call and let's grow your business with SEO, ads, and AI.

Book a Free Strategy Call →