← All articles
AI Visibility 101·27 September 2026·7 min read

What an AI visibility score actually measures

Search for an AI visibility score and most tools hand you one number out of 100. Here is what that number is usually hiding, and the separate readings we track instead.

By The Babel42 team

What an AI visibility score actually measures

You searched for "AI visibility score" because you wanted a number: a quick way to tell if your brand is doing well or badly when buyers ask ChatGPT or Gemini for a recommendation. Most tools that use the phrase will give you exactly that, a composite score out of 100. What that single figure hides is that it is usually built from several different readings that move independently, and two brands can land on the same score for opposite reasons.

What is an AI visibility score?

An AI visibility score is a summary of how often and how favourably an AI assistant mentions your brand when buyers ask for a recommendation in your category. Most tools that publish one compress it into a single figure, often out of 100, blending things like mention frequency, sentiment and share of voice into one number.

The trouble with a blended number is the one this whole idea runs into everywhere else: a category with two credible options and a category with twenty do not compare on the same scale, and a composite hides which one you're in. A strong appearance rate in a crowded field with a dozen credible rivals is a harder result to earn than the same appearance rate in a field with one obvious incumbent, and a single score can't tell you which situation you're looking at.

The six numbers a single score compresses into

Rather than roll results into one figure, we track six separate readings for every brand we measure, each with its own definition rather than folded into an index:

MetricWhat it answers
Appearance rateShare of journeys where your brand is named at all, whether or not it's recommended
AI win rateShare of journeys where the buyer's final pick was your brand
Weighted win rateWin rate after discounting wins the buyer had to be coaxed into
Shortlist rateShare of journeys where your brand reached the buyer's final two or three candidates
Share of AI voiceYour brand's mentions as a share of all brand mentions in the journeys measured
Own-site citation shareShare of the assistant's cited sources that are pages on your own domain

Read individually, each answers a different question. Appearance rate tells you whether the AI knows your category includes you at all. Win rate tells you whether it actually sends the buyer your way. Own-site citation share tells you whether the assistant is reading your own account of the facts or a summary written by someone else. A brand can score well on one and badly on another, and averaging them into a single figure erases exactly the gap that would tell you what to fix.

What this looks like on a real brand

Here's what those separate readings look like on a real dashboard rather than in the abstract. This is a demo workspace in Babel42's AI Visibility product, tracking an email marketing brand across two AI assistants over 16 buyer conversations.

Babel42's AI Visibility dashboard for a demo workspace, showing five separate number tiles: appearance 100%, AI win rate 25%, shortlist rate 69%, share of AI voice by mentions 24%, and share of AI voice by citations 10%, alongside a share-of-voice donut chart, a brand sentiment donut and a performance-by-AI-model bar chart

In this demo account, the brand is named in 100% of the 16 journeys we tracked, a strong appearance rate by any reading. Its shortlist rate is 69%, so most of the time it survives to the buyer's final two or three options. But its AI win rate is only 25%: it's in the room for almost every conversation and on the final shortlist most of the time, and still loses the actual recommendation three times out of four. Share of AI voice by mentions sits at 24%, worth a look but not dominant, and share by citations is lower still at 10%, meaning the assistant's cited sources lean even further away from this brand's own pages than the mention count alone suggests.

A single composite score would have to average those five numbers into one figure. Whatever number came out, it would tell you nothing about which of appearance, shortlisting or winning is the actual problem, and fixing the wrong one wastes the time you spent on it.

Why we don't compute one number out of 100

Leaving results as six separate numbers is a deliberate choice, not an oversight. Babel42's own metrics page states the reasoning plainly: "Babel42 does not roll your results into a single score out of a hundred, because categories differ enormously in how many credible brands an assistant has to choose from and a composite would hide that." A crowded category and a two-horse race need different benchmarks, and a single number applied to both is comparing things that aren't comparable.

In place of one score, the same page names the readings worth watching instead: "the gap between your appearance rate and your win rate, the gap between your win rate and your weighted win rate, and the trend in all three against the same set of questions." Those gaps are where the diagnosis lives.

The two gaps worth watching

A wide gap between appearance rate and win rate means the AI knows you exist and is still choosing someone else. That's a positioning problem: buyers reach the conversation, hear your name, and pick a competitor anyway, which usually means the assistant has more to say in favour of that competitor's fit for the specific need than for yours.

A wide gap between win rate and weighted win rate is a different problem. Weighted win rate discounts a win depending on how hard the buyer had to work to get there: a win counts in full when you're named unprompted in the assistant's first answer, and counts for less when you only turn up after the buyer pushes for more options. A brand can post a respectable raw win rate that's mostly made up of these harder-won recommendations, which is a more fragile position than the same win rate built from unprompted picks, because it depends on the buyer asking the right follow-up question rather than the assistant reaching for you first.

How do you check your own numbers?

You don't need a subscription to get a rough first read. Open ChatGPT, Claude and Perplexity in separate tabs and ask each one a genuine buyer question from your category, phrased the way a real prospect would ask it rather than a query built around your own brand name. Read to the end of each answer: note whether you're mentioned at all, whether you make it into a shortlist of options, and whether you're the one actually recommended at the close. Doing this across a handful of realistic questions, not just one, is what turns a single lucky or unlucky answer into a real reading; our fuller measurement method covers how to run it as a repeatable process rather than a one-off check.

Babel42's AI Visibility product runs this automatically: structured AI Buyers work through real, multi-turn shopping journeys across the assistants your buyers use, and appearance rate, win rate, weighted win rate, shortlist rate, share of AI voice and own-site citation share come back as tracked numbers rather than something you'd have to read off a transcript by hand. Each one has its own definition page if you want the full detail behind any of the six. The free plan runs one AI Buyer across any two of the seven assistants on a weekly cadence, no card required, enough to see your own version of the appearance-versus-win gap above before deciding whether to dig further.

Read your numbers as a set, not a score

The instinct to want one number is reasonable. A single score is easy to put in a slide and easy to compare month over month. But an AI visibility score that hides the gap between being named and being chosen is telling you less than it looks like it is. The readings worth tracking are the ones that show you where in the journey you're losing ground, appearance, shortlist or the final pick, because that's the one that tells you what to go and fix.

If you're new to the whole idea, what AI search visibility measures is the place to start. If you already track appearance and want to know whether the assistant is reading your own pages or someone else's when it cites a source, what AI citation rate measures covers that gap on its own terms. And once you know where the gap is, our playbook for improving AI search visibility covers what to do about it.

Enjoyed this?

Get the next dispatch in your inbox. No spam, unsubscribe anytime.

Occasional dispatches on listening, trends and the Babel42 roadmap. No spam, unsubscribe anytime.

Start listening free