← All articles
AI Visibility 101·7 September 2026·6 min read

There is no such thing as ranking in AI search. You rank in one assistant at a time.

We put the same buying question to ChatGPT, Claude, Gemini, Perplexity and Grok at the same moment, 55 times over. In 53 of those comparisons at least one assistant recommended a different brand. All five agreed twice.

By The Babel42 team

There is no such thing as ranking in AI search. You rank in one assistant at a time.

Five assistants, five different answers

Babel42 puts a realistic buying question to several AI assistants at once and records which brand each one ends up recommending. Same buyer, same wording, same moment. We went back over 55 of those comparisons across ChatGPT, Claude, Gemini, Perplexity and Grok to see how often they land on the same answer.

In 53 of the 55, at least one assistant recommended a different brand from the others. That is 96%. All five agreeing on a single brand happened twice.

The comparison only works because the question is genuinely identical. Every assistant in a comparison received the same opening question, written once and sent to all of them together, so nothing in the wording can explain the split. What is left is the assistants themselves.

That single number changes what "doing well in AI search" can even mean. A brand is not visible to AI. A brand is visible to one assistant, in one category, on one kind of question, and that tells you very little about the other four.

No two assistants agree with each other

The headline figure counts a comparison as a disagreement if any one assistant breaks from the rest, which is a low bar with five of them in the room. So we also measured every pair on its own: out of the comparisons where both members of a pair named a brand, how often was it the same brand?

PairAgreed on the same brand
Claude and Gemini44%
Claude and Grok38%
Gemini and Perplexity30%
Gemini and Grok25%
ChatGPT and Grok14%
Grok and Perplexity12%
ChatGPT and Claude11%
ChatGPT and Gemini10%
ChatGPT and Perplexity9%
Claude and Perplexity9%

The friendliest pair on the board agrees less than half the time, and six of the ten pairs agree less than 15% of the time.

Most teams check their brand in whichever assistant they personally use. On this evidence, that check tells you almost nothing about the rest. Being the answer in ChatGPT gives you roughly a one in ten chance of being the answer in Perplexity.

They are reading different webs

The reason shows up in the sources. Babel42 records which websites each assistant leaned on while it was making its recommendation, so we tallied the most-cited website for each one across every answer we captured.

AssistantMost-cited sourceThen
PerplexityRedditLinkedIn, Yelp
ChatGPTRedditWikipedia, TechRadar, Forbes
Google AI ModeLinkedIngoogle.com, Instagram
Google AI OverviewsYouTubeFacebook, Reddit
Geminivendor websitesCapterra

The five are not ranking the same web in a different order. They are reading different webs. Two of them are built mostly on what people say to each other on Reddit. One is built on LinkedIn profiles. One leans on video. One leans on vendors describing themselves.

Once you see the sources, the disagreement stops looking like a fault and starts looking like the obvious result. Two assistants reading two different halves of the internet were never going to arrive at the same shortlist.

What this means for getting recommended

Published research already tells us a good deal about what makes a page quotable. The 2024 KDD paper GEO: Generative Engine Optimization tested nine edits to real pages and found that adding statistics, quotations and cited sources were among the strongest ways to lift visibility in a generated answer. We measured the same thing across 164 sites and found most have almost nothing an assistant can quote.

What this comparison adds is the second half of the job. Being quotable decides whether a page can be used. Being present in the right sources decides whether the assistant ever reaches your page at all, and the right sources are different for each one.

In practice that means three things:

  • A well-argued Reddit thread can carry weight with ChatGPT and Perplexity, the two that cited Reddit above everything else, and do nothing at all for Google AI Mode.
  • A strong LinkedIn presence can carry weight with AI Mode, which cited LinkedIn above everything else, and very little with Perplexity.
  • Gemini was the only one of the five whose most-cited sources were vendor websites. For the other four the top source was somewhere you do not own, so your own site being excellent is necessary but not sufficient.

The practical order of work is to find which assistants your buyers actually use, find what those specific assistants read, and go and earn a place in those sources. Spreading the same effort evenly across all of them wastes most of it.

The limits of this measurement

Every number above comes from 55 paired comparisons, and we would rather state the limits than have them pointed out.

The sample is small, but it is paired. Each comparison is its own control: same buyer, same question, same moment, so the only variable is the assistant. That is why a 96% split is meaningful at this size. It is not a market survey and we would not present it as one.

A further 51 comparisons were set aside because at least one assistant in them did not settle on a single brand. Some assistants readily name one recommendation and others prefer to give a list. If every one of those 51 had quietly agreed, the disagreement rate would still be 50%.

Google AI Overviews and AI Mode are excluded from the agreement figures. Both answer in one pass rather than working through a decision, so they have no single recommendation to agree or disagree with. They appear in the source table only. The source table itself covers five of the seven.

The categories are whatever our customers happen to sell in, so nothing here supports a claim about any particular industry. It is a finding about assistants, not about markets. Brands and categories are not named, because they are real companies' competitive sets.

Where your own brand stands

The figures above describe assistants in general. The number that matters to you is your own, and it will differ by assistant in the same way.

Babel42 runs this comparison for your brand: the same buying question put to seven AI assistants, repeated on a schedule, with the winning brand and the sources behind each answer recorded every time. You see which assistants recommend you, which ones recommend a rival, and what each of them was reading when it decided.

See how AI assistants answer for your brand, or run the free audit to check whether your website gives them anything worth quoting in the first place.

Enjoyed this?

Get the next dispatch in your inbox. No spam, unsubscribe anytime.

Occasional dispatches on listening, trends and the Babel42 roadmap. No spam, unsubscribe anytime.

Start listening free