← All articles
AI Visibility 101·11 September 2026·7 min read

Does ChatGPT give different answers to the same question?

We asked AI assistants the same buying question twice, in the same run. The same assistant named the same winning brand again only 22% of the time, from 43% for Gemini down to 14% for Claude.

By The Babel42 team

Does ChatGPT give different answers to the same question?

A brand checks its name in ChatGPT once, sees itself recommended, and treats that as settled. Babel42's AI Buyer engine repeats a shopping journey inside the same run: the identical buying question, asked again, purely to see whether the second answer matches the first. We sampled 387 AI shopping journeys across seven assistants between 21 July and 7 September 2026, pulled out every case where an assistant named a winner on both tries, and counted how often the winner was the same brand.

Does ChatGPT give different answers to the same question?

Yes, most of the time. Across 114 repeat cases, where the AI Buyer engine asked an assistant the identical buying question twice within one run and it named a winner on both tries, the same brand won again in only 25 of them. That's 22%. Nothing about the question changed between the two asks: same product category, same buyer persona, same wording, same moment.

The other 78% of the time, the assistant landed on a different brand the second time round, with no new information and no new prompt to explain the switch. An AI recommendation, on this evidence, behaves less like a ranking and more like a draw that gets re-drawn on every ask.

How consistent each assistant is with itself

Consistency varies by assistant, and so does the sample size behind each row, so read the smaller ones as directional rather than settled.

AssistantRepeat casesSame winner both times
Gemini146 (43%)
ChatGPT389 (24%)
Perplexity387 (18%)
Claude213 (14%)
Grok30

Gemini looks the most self-consistent of the five, but 14 repeat cases is a thin base to build a claim on, and Grok's three cases are too few to read as anything at all. ChatGPT and Perplexity carry the largest samples here, and both landed on a different brand more often than they repeated one.

What to do instead of checking once

Treat a single named-or-not result as one draw, not a verdict. If you see your brand recommended in ChatGPT this morning and named the competition this afternoon, both readings are true and neither one is the whole picture; on this data, that swing is closer to normal than it is to a fault.

The practical fix is to ask the same kind of question more than once before drawing a conclusion, and to keep asking it on a schedule rather than once and done. A brand that only checks when someone in the team happens to think of it will see whichever draw happened to land that day and mistake it for the trend. Three or four repeats spread over a week say more than one lucky or unlucky check ever will, and the assistant-level table above shows why: even the more consistent assistants here still switched winners more than half the time. The same logic applies to a good result as much as a bad one: one win doesn't confirm you've arrived any more than one loss confirms you've fallen behind.

This is a different finding from assistants disagreeing with each other

Babel42 has also measured how often five AI assistants agree with each other on the identical buying question, in a separate piece. On the current corpus, they name the same brand in only 1 of 61 decidable comparisons, a 98% disagreement rate, with an honest floor of 53% even if every comparison an assistant sat out had gone the other way.

That's a related problem, but not the same one, and this one is harder to explain away. Two assistants naming different brands can be put down to two different products built two different ways. The same assistant naming a different brand a few minutes after it named another, on the identical question, has no such explanation on hand. Cross-assistant disagreement says assistants are not interchangeable. Same-assistant disagreement says a single assistant, on its own, is not settled either.

What one AI visibility check tells you

Not much, on this evidence. If an assistant only agrees with itself 22% of the time when asked twice in the same sitting, a brand that gets named once and stops checking has learned about one draw, not a ranking. That's an uncomfortable thing to publish as a company that sells AI visibility monitoring, and we're publishing it because it's what the data shows: a single check is close to worthless, and only a rate measured across many runs, over time, says anything reliable about whether an assistant tends to recommend you.

What a rate looks like instead of a single check

This is why the AI Buyer engine reports a rate rather than a yes-or-no. In the dashboard below, a demo brand tracked across two AI assistants appears in 100% of its recorded buying journeys but wins the recommendation in 25% of them, with a named rival as the usual winner. Neither the 100% nor the 25% comes from one conversation; both are built by repeating the same kind of question over many separate AI Buyer runs and counting the outcomes, the same method behind the table above.

AI Visibility overview dashboard showing a demo brand's appearance rate, win rate and share of AI voice

A brand watching only the first figure, appearance, would see itself named every time and conclude it's doing well. The win rate, tracked the same way over the same repeated checks, is the number that would have caught the gap between being mentioned and actually being recommended.

The limits of this measurement

Every figure above comes from repeat cases inside the same AI Buyer run: same buyer, same day, same wording, so this measures consistency within a sitting, not stability across weeks. We would rather state that limit than have it pointed out.

The 114 repeat cases only include instances where the assistant named a single winner on both tries; cases where it gave no clear winner, or offered a list instead of a pick, sit outside this count. Gemini's 14 cases and Grok's 3 are small enough that a couple of different draws would move the percentage a long way, which is why the assistant-level table above is read as directional, not as a settled ranking of who is most consistent. And the categories behind all 387 journeys are our own customers' mix of B2B software, legal services, consumer tools and the rest, not a claim about any one industry.

Two of the seven assistants Babel42 tracks, Google's AI Overviews and AI Mode, don't appear in either table. Both answer in a single pass rather than working through a back-and-forth, so there's no second draw to repeat and no winner to compare against a first one; that's how those two surfaces behave by design, not a gap in what we measured. The 114 repeat cases and the assistant table both come only from the five conversational assistants that can be asked the same thing twice.

Where your own brand stands

The figures above describe assistants in general. What matters for a specific brand is its own rate, checked the same way, repeatedly.

Babel42's AI Visibility product runs this AI Buyer comparison on a schedule rather than once: the free plan checks one AI Buyer against any two of the seven AI assistants it tracks every week, enough to see whether your own appearance rate and win rate move together or apart. What AI search visibility actually measures is the place to start if the appearance-versus-win distinction above is new, and the factors that earn a citation in the first place covers what to fix once you know your own numbers.

Enjoyed this?

Get the next dispatch in your inbox. No spam, unsubscribe anytime.

Occasional dispatches on listening, trends and the Babel42 roadmap. No spam, unsubscribe anytime.

Start listening free