Back
Experiment

What ChatGPT, Perplexity, Gemini & Claude actually cite

Mike Price·September 2026·8 min read

I asked four AI engines the same two questions and logged every source each one cited. They recommended the same handful of brands - and then backed those recommendations with almost entirely different evidence.

69unique domains cited across the study
93%cited by only one of the four engines
0domains cited by three or more engines
What I found

They agree on conclusions, not on sources. All four engines recommended NoGood, iPullRank and First Page Sage - yet across both questions, only five domains were cited by more than one engine, and none by more than two.

Each engine has a distinct source appetite. Perplexity cited 35 domains (SEO media, PR/newswire, Reddit); ChatGPT cited just 8 (mostly official docs and research); Gemini and Claude leaned on small niche blogs - different ones.

So "get cited by AI" is really "get cited across several different source ecosystems at once."

The setupHow I ran it.

On 30 September 2026 I sent the same two prompts to four engines, each with live web search enabled, via their APIs:

  • Q1 (commercial): "What are the best answer engine optimization (AEO) agencies or consultants?"
  • Q2 (informational): "How can a company get its brand cited by AI search engines like ChatGPT and Perplexity?"

The models were ChatGPT (gpt-5.6-sol), Perplexity (sonar-pro), Gemini (gemini-3.8-flash) and Claude (claude-sonnet-5). For each answer I captured the full list of cited source URLs, then reduced them to domains. It's a small sample - two questions, one run, one day - so treat it as a snapshot, not a census. But the pattern was stark enough to be worth showing.

Finding 1They agree on the brands.

On the commercial question, the four engines converged hard. Three names showed up in every single answer, and a second tier appeared in two of the four:

RecommendedGPTPlxGemCld
NoGood✓✓✓✓
iPullRank✓✓✓✓
First Page Sage✓✓✓✓
Omniscient Digital✓✓––
Siege Media✓✓––
Minuttia––✓✓

If you only looked at the recommendations, you'd conclude these engines basically agree. That conclusion falls apart the moment you look at how they got there.

Finding 2They disagree on the sources - almost completely.

Across both questions the four engines cited 69 unique domains. Of those, 64 (93%) appeared in only one engine's answers. Just five domains were cited by two engines - firstpagesage.com, minuttia.com, nogood.io, reddit.com, yesoptimist.com - and not a single domain was cited by three or more.

On the "how to get cited" question specifically, the overlap was even starker: the four engines shared zero sources. Four answers to the same question, no common ground in the citations at all.

The engines reached the same conclusions from four almost completely separate slices of the web.

Finding 3Each engine has a source personality.

The counts and the character of the sources were very different:

EngineDomainsWhat it leaned on
ChatGPT8The fewest, and the most authoritative: official docs (OpenAI, Bing, Google), academic research (arXiv), plus a few niche AEO blogs
Perplexity35By far the most, and the widest range: SEO media (Search Engine Land, Moz, Ahrefs, Semrush, HubSpot), PR/newswire (Forbes, USA Today), and Reddit
Gemini18Mostly small niche AEO/GEO vendor blogs, plus Reddit and YouTube, surfaced through Google's grounding layer
Claude13A different set of small niche blogs, plus community and professional sources like Quora and LinkedIn

ChatGPT was the most conservative - when asked how to get cited, it went almost entirely to primary sources (OpenAI's own crawler docs, Bing Webmaster Tools, Google's structured-data docs). Perplexity was the opposite: broad, current, and happy to cite newswire and Reddit. Gemini and Claude both leaned on small, niche industry blogs - but with virtually no overlap between their picks.

Finding 4The fan-out is real, and it compounds.

Two of the engines exposed their query fan-out - the multiple sub-searches they run behind a single question. For Q1, ChatGPT fired eight sub-queries (including targeted site: searches against specific agencies); Gemini ran two. Claude's own answer described the mechanism plainly: when a brand shows up across several of those sub-queries, a process it called Reciprocal Rank Fusion compounds that repetition into a citation.

That's the real lever. You're not trying to rank once - you're trying to appear across many related sub-queries so the fusion step keeps surfacing you.

What it meansPractical takeaways.

  • Don't optimize for one engine. With 93% of sources unique to a single engine, winning ChatGPT tells you almost nothing about Perplexity or Claude.
  • Diversify where you appear. Own-site content, SEO media coverage, PR/newswire, Reddit and niche industry blogs each feed different engines. You need presence across the whole spread.
  • Reddit and listicles matter. They were among the very few sources multiple engines shared - and the "best agencies" roundups clearly shaped every engine's recommendations.
  • Consensus wins the recommendation. The brands named by all four are the ones repeated across many independent roundups. Being everywhere is what gets you into the answer.
  • Aim to appear across sub-queries, not just to rank for the head term - that's what the fan-out rewards.
Caveats. This is a deliberately small experiment: two questions, one run, a single day, via APIs with web search on - which approximates but isn't identical to the consumer apps. LLM answers are non-deterministic and vary by wording, location and time. Treat the exact numbers as a snapshot. The pattern - convergent recommendations, divergent sources - is the part I'd expect to hold, and it lines up with the mechanism: each engine reads a different index. I wrote about that foundation in ranking factors across the major AI engines.
Work with me

I help brands get visible across AI search - measuring citations per engine, then building the content, structure and off-site presence each one rewards. From audit through to implementation.