How to choose an AI visibility tool
Judge an AI visibility tool on its measurement, not its interface. Ask how many assistants it covers, how many times it samples each question, whether the questions are phrased the way buyers ask them, who you are compared against, and whether results carry confidence intervals. A tool that cannot answer those is selling a dashboard, not a measurement.
Updated 19 July 2026
Start with the job, not the vendor
Most disappointment in this category comes from buying the wrong kind of product. Before comparing anyone, decide which of these you actually need:
- Diagnosis — why are we invisible? A readiness audit answers this, and the technical half is cheap.
- Awareness — are we being named, and is it changing? That is monitoring.
- Evidence — a number we can defend externally, against a named field. That is an index.
- Execution — someone to do the work. That is an agency.
The tools survey groups the market that way. Everything below assumes you want the second or third.
The six questions
1. How many assistants, and which?
ChatGPT, Gemini, Perplexity and Google's AI Overviews disagree constantly. A tool measuring one and reporting "your AI visibility" is describing one weather station and calling it the climate.
Ask which providers, and whether the answers are grounded — able to search the live web — or purely from training. Those produce genuinely different results, and a tool that does not distinguish them cannot tell you which lever to pull.
2. How many samples per question?
This is the question that most cleanly separates serious tools from demos.
Assistants are non-deterministic. Ask "best cigar shop in the UK" five times and you get overlapping but different brands. One answer is one draw from a distribution, and a percentage computed from a single draw per question is not a rate — it is a coin flip with a decimal point.
Ask how many samples per prompt. If the answer is one, treat every number as indicative at best.
3. Are the questions phrased the way buyers ask them?
There is a large difference between:
"Tell me about Acme Cigars" — will always produce something flattering
"Where can I buy Cuban cigars online in the UK?" — the question that decides whether you exist
Ask to see the actual prompt set. If it is mostly brand-name lookups, the tool is measuring whether the model has heard of you, not whether it recommends you.
4. Who are you compared against?
Share of voice is a fraction, and the denominator is a decision somebody made.
If the comparison set omits real competitors, your share is overstated — and you will never know, because the missing names simply do not appear. We found exactly this in our own data: a category we had measured for weeks turned out to have a cohort accounting for under half of all the brands assistants actually named, which meant every share figure in it, including our own customer's, was inflated.
Ask: who is in the set, how was it chosen, and can I see it? A vendor who cannot show you the denominator is showing you a number they cannot explain.
5. Do the results carry confidence intervals?
A mention rate of 12% from 30 samples and 12% from 600 samples are very different claims. Without an interval you cannot tell whether this month's move is progress or noise — and you will spend real money reacting to noise.
6. What happens when the methodology changes?
The most dangerous chart in this category is one where the number moved because the tool changed — new questions, new judging logic, a new provider added.
Ask how they version the question set, and what happens to your history when it changes. The honest answer is that a trend across a methodology change is not a trend, and it should be broken rather than quietly redrawn. If a vendor claims continuous comparable history while also improving their method, one of those two claims is false.
Red flags
- Guaranteed mentions or placements. There is nothing to buy. This is the clearest signal that someone does not understand the product they are selling.
- A single headline score with no method published. If you cannot see how it is computed, you cannot defend it to anyone who asks.
- Rankings with no stated field. See question 4.
- "We optimise your llms.txt" as a headline offer. See our llms.txt guide — it is unproven, and leading with it suggests a thin bench.
- Screenshots as evidence. One conversation proves nothing about a distribution.
The questions we find uncomfortable
A buyer's guide written by a vendor is worth very little unless it holds the vendor to the same standard, so:
- Our corpora are small and deliberately curated — tens of questions, not hundreds. We think a small, versioned, defensible question set beats a large unexamined one, but if you need breadth of coverage, that is a real limitation and you should weigh it.
- We publish periodic readings, not live monitoring. If you want daily alerting, we are not the right shape.
- Our own cohorts have been incomplete, and we found it by measuring rather than by being told. We publish the methodology precisely so this is checkable rather than a matter of trust.
- We cannot tell you what a model will say tomorrow. Nobody can. A reading is a repeatable measurement, not a forecast.
If a competing vendor will not write the equivalent list about themselves, that tells you something worth knowing.