How do you calculate a brand visibility score?
A brand visibility score is the share of AI answers about your category that name your brand, usually expressed as a percentage or on a 0–100 scale. The formula is simple: mentions divided by total answers. What makes a score credible or worthless is decided before the division — by the denominator, the sample size, and what you allowed to count as a mention.
Updated 26 July 2026
Every vendor publishing this metric uses roughly the same formula, and it is not a secret:
brand visibility score = (answers naming your brand ÷ total answers) × 100
Ask 100 buyer-phrased questions across the assistants your customers use, count the answers that name you, divide. If you appear in 22, you score 22.
The arithmetic is trivial. Everything that decides whether the number is worth anything happens before it — in four choices that are rarely published.
1. What counts as "total answers"
The denominator sets the score. Ask ten questions where your brand is the obvious answer and you will score beautifully; ask the questions your buyers actually type and you may not. A score built on a question set chosen by the brand being measured is a self-assessment.
This gets worse when the score is expressed as share against competitors rather than against total answers. When we first measured UK cigar retail, the field we had defined captured under half the merchant mentions the assistants actually made — the rest were businesses nobody had thought to put on the list. Every share figure we could have shown at that point was inflated roughly two-to-one, not through any error in the counting, but because the list was short. Share of voice goes through what fixed it.
2. How many times you asked
Assistants are non-deterministic. The same question, asked twice, returns different brands. So a visibility score is an estimate of a rate from a sample, and it carries a confidence interval whether or not anyone shows you one.
A real example from our readings: a mention rate of 10.7%, with a true value somewhere between 7.5% and 15.1%, across 252 sampled answers. That is a properly sized sample and the range is still six points wide.
Two consequences follow, and both are unpopular:
- A bare percentage is a midpoint. Anyone quoting "you scored 22" without a range has an interval they have not shown you.
- Movement inside the range is not progress. A score drifting 8 → 9 → 8 is dice. It will be sold to you as improvement anyway.
3. What you allowed to count as a mention
A string match is not a mention. Brands with generic or common-word names get credited with text that has nothing to do with them — we hit this with our own name, whose first measurements counted appearances of a common phrase as appearances of us. Separating an attributable mention from a coincidental phrase match is a filter someone has to build, and a score computed without one is inflated in a way that grows with how ordinary your name is.
The reverse also happens: brands named by a shortened or misspelled form go uncounted.
4. Whether you blended mention, recommendation and citation
These are three different outcomes and they fail independently:
- Named — your brand appears in the answer.
- Recommended — the answer puts you forward as a suggestion, not a passing reference or a warning.
- Cited — your own page is linked as a source.
A composite that averages them produces a middling number that hides which one you are failing, and they have different fixes. Being cited but never recommended is a positioning problem; being recommended but never cited is usually fine.
What a defensible score looks like
Four properties, all of which you can check on any vendor's score before buying:
- The question set is published and fixed, so the score cannot be improved by changing the exam.
- The sample size and interval are shown, so you can tell a result from noise.
- It decomposes. A score you cannot open up into its inputs is an opinion with a decimal point.
- The method is versioned. When the question set or the set of assistants changes, older readings stop being comparable. A serious score breaks its own trend line and says so — when we added a fourth assistant, readings taken before and after became non-comparable on purpose.
Calculating one yourself
You can do this by hand, and it is a good idea before you pay anyone. Write 15–20 questions a buyer would actually ask, with no brand names in them. Ask each one three times, in fresh sessions, across at least three assistants. Record every brand named in every answer — not just yours. Divide.
That gives you a defensible score for a day's work. The parts that do not scale by hand are re-running it identically each quarter, and building the field of competitors from what the assistants named rather than from what you assumed.
Our own score works slightly differently — it weights whether assistants name you at 60% and whether they can read your site at 40%, so that the two failure modes stay visible. What the AVI Score is made of covers that, and the methodology states the rules in full.