How does ChatGPT choose its sources? In search mode it rewrites your question into targeted queries, sends them to partner search indexes (OpenAI names Microsoft and Shopify), ranks the results on relevance and reliability, and cites the pages that survive. Perplexity and Google AI Overviews do the same job with 3 different pipelines, which is why a page can be cited by one and invisible to the other 2.
A Shopify store owner posted this in the Shopify Community last year: "Why does ChatGPT cite Reddit and gift guides for my products instead of my actual Shopify pages?" That one sentence is the whole problem. The engines are not ignoring you. They are picking sources using rules you have never seen written down.
This guide lays those rules out, engine by engine, using each provider's own documentation. No guessing, no recycled LinkedIn theories.
In this guide
- The short answer: 3 engines, 3 source pipelines
- How does ChatGPT choose its sources?
- How Perplexity picks what to cite
- Where does Google AI Overview get its information?
- Side by side: how the 3 engines choose sources
- What this means for a DTC brand, a coach, and a home service business
- What to fix first, in order
- What I am not sure about
- The verdict
The short answer: 3 engines, 3 source pipelines
ChatGPT, Perplexity, and Google AI Overviews all end in the same place, a paragraph with links under it. They get there 3 different ways.
ChatGPT search leans on partner search indexes and a ranking layer OpenAI keeps private. Perplexity runs its own live retrieval first and hands the results to a language model second. Google AI Overviews fan your query out into many related searches against Google's own index, then pick supporting pages from the results.
Different pipelines mean different winners. If you treat "AI search" as one thing and optimize once, you are optimizing for none of them. That is the mistake behind most answer engine optimization advice I read.
How does ChatGPT choose its sources?
ChatGPT has 2 different ways of answering, and they source information completely differently. Most confusion about ChatGPT sources comes from not knowing which mode you are looking at.
Two modes, two different answers
Mode 1 is the model answering from training data. There are no live sources here. When ChatGPT answers this way, it is recalling patterns from text it was trained on, which is why it can name a brand it read about in 2024 and miss a competitor that launched in 2026. Ask it for a source in this mode and it will often invent one. That is where "chatgpt fake sources" and "chatgpt hallucinating sources" come from, both real Google autocomplete phrases.
Mode 2 is ChatGPT search. This is the mode with real citations, and it is the only mode you can influence with your website. When search runs, ChatGPT goes out to the live web through partner indexes and cites what comes back.
What OpenAI actually says
OpenAI's ChatGPT search help article is short, but it gives away the mechanism. It says ChatGPT search "sometimes partners with other search providers" and "typically rewrites your query into one or more targeted queries that it sends those providers." Microsoft and Shopify are the 2 partners named.
On ranking, the exact line is: "ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed."
On eligibility, OpenAI says to "allow OAI-SearchBot to crawl the site" and confirm your host or CDN allows traffic from its published searchbot IP addresses. OAI-SearchBot is the crawler for search results and links. It is separate from GPTBot, which OpenAI uses for training. Blocking GPTBot to protect your content while leaving OAI-SearchBot open is a real, supported choice.
Read those 3 statements together and the picture is clear. If you are not indexed and ranking in Bing for the query, you are not in the candidate pool for ChatGPT search. If you are a Shopify merchant, your product data has a second door in through the Shopify partnership.
Why Reddit and listicles keep winning
Back to the Shopify owner's question. Reddit threads and "best gifts" listicles get cited for product queries because the retrieval step is looking for pages that answer the user's actual question. The user did not ask "tell me about Brand X." They asked "what is a good gift for a runner." A Reddit thread with 40 people comparing options answers that. A product page with 3 bullet points and a price does not.
The ranking layer also favors what OpenAI calls reliable sources. In practice that means pages with third-party consensus, recognizable domains, and content that reads like an answer rather than an ad.
What a small brand can influence
You cannot buy placement. You can control 4 things: being crawlable by OAI-SearchBot, ranking in Bing for the questions your buyers ask, having pages that answer those questions in plain language in the first 100 words, and getting mentioned on the third-party pages ChatGPT already trusts. I show the page-level rewrite in How to Get Cited by ChatGPT, with a real before and after.
How Perplexity picks what to cite
Perplexity is the most transparent of the 3 about the fact that it searches, and the least transparent about how it ranks.
Retrieval first, model second
Perplexity's own help center describes the order of operations: it "searches the internet, gathering information from authoritative sources like articles, websites, and journals," then uses language models (it names GPT-5 and Claude 4.6 Sonnet) to understand the query and write the answer. Every answer "includes numbered citations linking to the original sources."
Retrieval first means the model never sees your page unless retrieval found it. Page structure, freshness, and a clear match between your heading and the user's question decide whether you enter the pool. Writing quality decides whether you get quoted once you are in it.
PerplexityBot vs Perplexity-User
Perplexity runs 2 different fetchers, and its bot documentation spells out the difference. PerplexityBot is "designed to surface and link websites in search results on Perplexity." It respects robots.txt and Perplexity recommends allowing it. Perplexity-User "supports user actions within Perplexity." When a user asks a question, it "might visit a web page to help provide an accurate answer and include a link to the page," and the documentation states plainly that "this fetcher generally ignores robots.txt rules."
That second fetcher is why blocking Perplexity in robots.txt does not reliably keep you out of Perplexity answers. It is also why Cloudflare publicly reported in August 2025 that it had observed Perplexity using undeclared crawlers to get around no-crawl directives. Whatever you think of that behavior, it changes the strategy: you cannot opt out cleanly, so you are better off making sure the page it fetches is the right one.
Where Perplexity differs from ChatGPT
3 practical differences. Perplexity cites more sources per answer, typically 5 to 10 numbered citations against ChatGPT's smaller set. Perplexity weights recency harder because it fetches live for every query, so a dated "Last updated" line and current-year facts matter more here. And Perplexity is less dependent on Bing rankings, because it runs its own retrieval, so a page that is buried on Bing still has a shot if it matches the question tightly.
Where does Google AI Overview get its information?
From Google's index. That is the whole answer, and Google says so in its own documentation. The interesting part is how it picks from that index.
Query fan-out
Google's Search Central page on AI features says both AI Overviews and AI Mode "may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources, to develop a response." It then says that while the response is being generated, "our advanced models identify more supporting web pages, allowing us to display a wider and more diverse set of helpful links associated with the response than with a classic web search."
Translate that: a user asks 1 question, Google silently runs 6 or 8 related searches, and the pages that rank for those sub-questions become the candidate pool. You can be cited in an AI Overview for a query you do not rank for, if you rank for one of the sub-questions Google fanned out to.
Why page 1 is necessary, not sufficient
Google states the only requirement: your page must be "indexed and eligible to be shown in Google Search with a snippet." So traditional SEO is the entry ticket. But the fan-out means the page that wins is the one that answers a specific sub-question cleanly, not the one that ranks number 1 for the broad head term. A 4,000-word guide that buries the answer in paragraph 12 loses to a 900-word page that answers it in the first line.
What Google says, and what it does not say
Google is explicit that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary," and that "there's also no special schema.org structured data that you need to add." Take that at face value. Schema will not get you in.
What Google does not say is how it chooses between 2 eligible pages. It does not publish the weighting. Anyone selling you a "guaranteed AI Overview placement" formula is selling you something Google itself says does not exist.
Side by side: how the 3 engines choose sources
| Question | ChatGPT search | Perplexity | Google AI Overviews |
|---|---|---|---|
| Where do candidate pages come from? | Partner search indexes (Microsoft, Shopify named) | Perplexity's own live retrieval | Google's own index via query fan-out |
| What gets you into the pool? | Crawlable by OAI-SearchBot, ranking in Bing | Tight match to the question, fresh content | Indexed and eligible for a Google snippet |
| Citations per answer | Fewer, often 3 to 5 | More, often 5 to 10 | Varies, shown as link cards |
| How much does recency matter? | Medium | High | Medium, depends on the query |
| Third-party vs first-party bias | Strong pull toward Reddit, reviews, listicles | Moderate, will cite brand pages that answer directly | Follows what already ranks |
| Can robots.txt keep you out? | Yes, block OAI-SearchBot | Not reliably, Perplexity-User ignores it | Yes, standard Google controls apply |
| What to fix first | Bing indexing and answer-first pages | Freshness dates and question-matched headings | Rank for the sub-questions, not just the head term |
What this means for a DTC brand, a coach, and a home service business
If you sell products online. Your product pages are losing to Reddit and gift guides because they do not answer buying questions. Add a real "who this is for and who it is not for" section, a comparison against the 2 alternatives your buyers actually consider, and a plain-English answer to the top 3 questions from your support inbox. If you are on Shopify, keep your product data clean, because that partnership is a second route into ChatGPT.
If you sell coaching or a course. Your offer page is a sales page, and none of the 3 engines cite sales pages. They cite pages that explain. Write the page that answers "how does [your method] work" in the first 80 words, with your name on it, and link it from the offer page. That is the page that gets cited. The offer page gets the click.
If you run an HVAC, plumbing, or electrical company. Google AI Overviews matter most to you, and they pull from what already ranks locally. Your Google Business Profile and your service pages are the candidate pool. The fan-out question is not "best plumber Phoenix," it is "why is my AC blowing warm air," and the local company whose page answers that is the one that gets cited.
What to fix first, in order
- Confirm all 3 crawlers can reach you. Check robots.txt and your CDN rules for OAI-SearchBot, PerplexityBot, and Googlebot. One blocked bot removes you from one engine entirely.
- Check you are indexed in Bing, not just Google. Bing Webmaster Tools takes 10 minutes to set up. Without Bing, ChatGPT search cannot find you.
- Rewrite your top 5 pages to answer in the first 100 words. Question as the heading, direct answer as the first paragraph, proof under it. Every engine's retrieval rewards this.
- Add a visible "Last updated" date and keep it honest. Perplexity in particular weights recency, and a dated page beats an undated one on the same topic.
- Rank for the sub-questions, not only the head term. Map the 6 to 8 questions Google would fan out to for your main keyword, and make sure one page on your site answers each.
- Get mentioned where the engines already look. A single real mention in a Reddit thread, an industry roundup, or a review site does more for ChatGPT citations than 10 new blog posts on your own domain.
If you want to see where you stand before doing any of this, run the AI search visibility audit first. It takes 20 minutes and it tells you which of the 3 engines is the real problem.
What I am not sure about
None of the 3 companies publishes a ranking specification. OpenAI says "multiple factors." Perplexity says "authoritative sources." Google says its models "identify more supporting web pages." The citation counts and recency weights in the table above come from my own testing across client and personal sites, not from a published formula, and they will drift as the products change.
I also cannot tell you how the Shopify partnership ranks one merchant above another inside ChatGPT. OpenAI has not published that. If someone tells you they know, ask them for the document.
The verdict
So, how does ChatGPT choose its sources? Through partner search indexes, a private ranking layer, and a strong preference for pages that answer the question rather than pitch a product. Perplexity does its own retrieval and rewards freshness. Google AI Overviews fan out your query and pull from pages that already rank for the pieces.
3 engines. 3 pipelines. 1 thing they all reward: a page that states the answer first and proves it second. Build that page and you have a shot at all 3. Skip it and no amount of schema, llms.txt files, or "AI optimization" will save you.
Want to know which of the 3 engines is actually ignoring you?
I run SEO, AEO, and AI search visibility work for US businesses on a free 30-minute strategy call. You bring the site, I show you where it drops out of each pipeline. No pitch deck.