How do you track product visibility across AI assistants at SKU scale?
You do not track it per SKU. The arithmetic does not work and the questions would be wrong — assistants answer at the level of a need, not a part number. Measure a curated set of buyer-phrased questions covering the product families that carry your revenue, and treat individual SKU visibility as an inference from that.
Updated 26 July 2026
This is the question a retail team actually asks: we have thousands of SKUs and we need to know which ones AI assistants recommend. The honest answer is that per-SKU tracking is the wrong shape of solution, and it fails for two independent reasons — either one of which is enough on its own.
Reason one: the arithmetic
A defensible visibility reading is not one query. It is a question set, asked repeatedly, across several assistants, because answers are non-deterministic and a single ask is a coin flip.
Our own brand-level readings run roughly 300 sampled answers across four providers, and cost a little under £2 per brand per pass at current provider pricing — the great majority of that being grounded-search fees rather than tokens.
Apply that per SKU:
- 1,000 SKUs ≈ £2,000 per pass.
- A weekly trend is 52 passes — six figures a year, to produce a dashboard.
- And that is the floor, because it assumes each SKU needs no more sampling than a whole brand does.
It needs more, not less. Which brings the second problem inside the first: a rare SKU has a near-zero surface rate, and low rates carry proportionally wider confidence intervals. You would spend the most money to produce the numbers you can trust least. A reading of "0.4%, somewhere between 0.0% and 2.1%" for a £12,000 spend is not an insight.
Reason two: the questions would be wrong
This one is fatal regardless of budget.
Buyers do not ask for SKUs. Nobody types "is product ref BX-4471-K available". They ask "waterproof walking boots for wide feet under £150" — a need, with constraints. The assistant answers at that level and returns three to five options.
So a per-SKU question set is a set of questions no buyer asks. You would be measuring your performance on a corpus you invented, precisely the failure the whole discipline keeps repeating: a number that is internally consistent, fully reproducible, and about nothing.
Measure at the level the assistant actually answers. That is the need, or the product family — not the part number and not, usually, the whole brand either.
What to measure instead
Pick the families that carry the revenue. Concentration is normally sharper here than in classic search, because an assistant returns three to five options rather than a page of forty. The long tail of your catalogue has less upside in AI answers than it does on a category page, so weighting by revenue rather than by SKU count costs you less than it feels like it should.
Write 15–25 buyer-phrased questions covering those families, containing no brand names and no product codes — constraints, budgets and use cases instead, the way a shopper phrases it.
Ask each across three or more assistants, repeatedly, and count: does a product of yours appear, which one, in what position, and who is offered as the place to buy. What to count in shopping answers goes through those four numbers.
Then infer downward. If your walking-boot family surfaces well and your insulated-jacket family does not, you have a finding that applies to hundreds of SKUs — and it is far more actionable than a per-SKU rate, because the fix is at family level too.
The lever that does scale to the whole catalogue
Here is the part that genuinely applies to all ten thousand items, and it is not tracking. It is legibility.
Platforms describe their product selection as weighing relevance, completeness of product data, clarity, competitive pricing and merchant trust signals. Every one of those is a property of your feed and your product pages, and every one is a fix-once-apply-to-all change:
- Complete structured attributes on every item — size, material, colour, compatibility, the constraints buyers actually filter on. Missing attributes make a product unmatchable to a constrained question, which is the only kind of question anyone asks.
- Titles that describe the thing, not internal nomenclature.
- Accurate price and availability, kept current. A stale feed gets you surfaced and then abandoned.
- Unambiguous product identity — consistent identifiers, so the same item is not treated as three different products.
- Server-rendered product pages. If the price and specification only appear after JavaScript runs, assume a crawler saw an empty page. This is checkable in minutes with a free scan.
Catalogue hygiene is unglamorous and it is the highest-leverage work available to a large retailer, precisely because it multiplies across every SKU without needing a single one of them tracked.
What this approach genuinely costs you
Stated plainly, because the trade is real: you will not know how any specific SKU is doing. If a single product is strategically critical — a hero item, a new launch, a licensing obligation — measure it deliberately, with its own question set, as if it were a brand. That is a defensible decision to spend, made for a handful of items. It is not a strategy for a thousand.
And whatever set you choose, re-run it the same way each time. Same questions, same assistants, same counting rules — change any of them and you have measured your method rather than your catalogue. The repeatable audit method sets out the procedure.