AThe AI Visibility Index
← Guides

How do you measure product visibility in AI shopping answers?

Measure discovery, not checkout. Count how often assistants surface your products in response to buyer questions, across enough samples to produce a rate with a range around it. The in-chat purchase mechanism has already been launched and withdrawn once; whether you are named at all is the durable thing to instrument.

Updated 26 July 2026

Shopping answers are not brand answers with products in them. They are a separate surface with a separate failure mode, and measuring them with a brand-mention metric will tell you comfortable things that are not true.

Three different events, routinely counted as one

When someone shops through an assistant, there are three distinct things that can go right or wrong, and they fail independently:

  1. Your brand is named in an answer about the category. This is ordinary brand visibility — how it is scored applies unchanged.
  2. A specific product of yours is surfaced — an actual item, with an image and a price, in a product-shaped answer. A brand can be well known to an assistant and have none of its products surface.
  3. You get the sale rather than a marketplace. Your product appears, and the route offered goes to a reseller instead of you. Your product won; your margin did not.

The third is the one retailers discover late, because a brand-mention dashboard records it as a success.

Do not anchor the metric to the checkout mechanism

This is the part worth being blunt about, because it has already caught people out.

OpenAI launched Instant Checkout — buying inside the chat — at the end of September 2025, and it was reported withdrawn on 4 March 2026, roughly five months later, after low merchant adoption. The direction since has been discovery-first: surface products, route the shopper to a merchant app or storefront to complete. Agentic-purchase work continues through the Agentic Commerce Protocol with payment partners, so the mechanism is likely to reappear in some form.

We are stating that as reported, from mainstream coverage in March 2026 rather than from our own measurement, and this page gets corrected if it changes again. Which is precisely the point:

A metric anchored to a mechanism dies with the mechanism. Anyone who built a 2025 reporting line around "in-chat checkout conversions" spent five months instrumenting something that no longer exists. The discovery layer — whether an assistant puts your product in front of a buyer — survived the change untouched, because it is the part that is actually about you rather than about the platform's current commercial experiment.

Measure discovery. Treat checkout as plumbing that will keep moving.

What to actually count

Four numbers, each of which needs a sample size and an interval rather than a single check:

Shopping-trigger rate. What fraction of your category's buyer questions produce a product-shaped answer at all, rather than a paragraph of prose? This varies enormously by category and it sets the ceiling on everything else. If only a fifth of questions in your category trigger a shopping-style response, product optimisation has a small denominator and brand visibility matters more.

Product surface rate. Of the answers that do go product-shaped, how often does one of yours appear?

Position within the set. Assistants return three to five options, not forty. Being fourth is materially different from being first in a way that a rank-20 position in classic search never was.

Merchant capture. When your product is surfaced, who is offered as the place to buy?

Each of those is a rate estimated from sampled answers, so each carries a confidence interval. In our own brand-level readings a mention rate of 10.7% came with a true range of 7.5–15.1% across 252 sampled answers — and product-level rates are lower and therefore wider, not narrower. Anyone quoting you a bare product visibility percentage is quoting the midpoint of a range they have not shown you.

"Which is better — an AI-visibility platform or an SEO suite?"

We get asked this in the specific form "Profound or Ahrefs". We are not going to answer it as a feature comparison, and it is worth saying why rather than dodging: we publish an index in this category, and our tools guide states the policy — we do not publish competitors' feature lists or prices, because this market changes monthly and a stale table presented as current is the exact artefact we spend our time arguing against.

The structural difference is durable and does answer the question underneath:

Tools that sample assistants directly — Profound and most AI-visibility natives — run a question set against the assistants on a schedule and record what comes back. You control the exam. The risk is that you write questions your buyers would never ask, and then measure your performance on an invented corpus.

Tools built on observed search demand — Ahrefs' Brand Radar is the clearest example — derive their prompts from real search behaviour. You get questions people demonstrably ask. The risk runs the other way: less control over the exact set, and search-derived prompts are not always the same shape as the conversational ones people type into an assistant.

So the choice is not better-versus-worse, it is which error you would rather make. If your worry is "we are optimising for questions nobody asks", take the demand-derived route. If your worry is "we need a fixed, versioned exam we can trend for years", take the controlled question set — and check that the vendor actually versions it, because a question set that changes silently makes every trend line fiction.

The fuller buying questions include the ones that are uncomfortable for us to answer.

Where we are the wrong choice: we measure at brand and category level and publish it as an index. We do not run product-level shopping monitoring, so if SKU-shaped dashboards are what you need, one of the above fits better than we do.

What nobody can tell you

Platforms describe their product picks as unsponsored and ranked by relevance, and state the signals they weigh — relevance, completeness of product data, clarity, competitive pricing, merchant trust. That is the stated position, and we have no way to verify it from outside. Any vendor claiming to know the ranking function is guessing with confidence.

What you can verify is your own catalogue's legibility, and at scale that is the lever that matters — tracking product visibility at SKU scale covers why.

Common questions

Is product visibility the same as brand visibility?
No, and treating them as one is the common mistake. Your brand can be well known to an assistant while none of your individual products surface in product-shaped answers, and your product can surface while the buying route offered goes to a marketplace rather than to you. They are three separate events that fail independently and need counting separately.
Can I still measure in-chat checkout?
There is much less to measure than there was. OpenAI's Instant Checkout was reported withdrawn in March 2026, about five months after launch, with the direction shifting to surfacing products and routing shoppers to merchant apps and storefronts. Agentic purchasing continues to be developed, so expect the mechanism to change again — which is the reason to instrument discovery rather than checkout.
How many samples do I need for a product visibility number?
More than for a brand number, because product-level rates are lower and low rates carry proportionally wider intervals. As a floor, treat anything under a few hundred sampled answers per question set as directional only, and refuse to read movement that sits inside the range.
Do AI shopping answers favour big retailers?
Not reliably, which is the recurring surprise. Platforms state that picks are ranked on product-data quality, relevance, price competitiveness and merchant trust rather than on size, and in the brand-level indexes we run, audience size predicts AI visibility poorly. A small retailer with complete, legible product data and genuine third-party coverage competes better here than in paid channels.

Read next

See how AI reads your brand

Get your AVI Score out of 100 — and a clear plan to rank higher in AI answers.